AI Adoption in Market Research: Past the Pilot
AI adoption in market research stalls after the pilot. What separates a pilot from production, what the public numbers show, and the gate a step has to cross...
An insights team has a tool that reads open ends, the answers respondents type in their own words. It has run on every tracker wave since March of last year, and the two analysts who use it would fight anyone who tried to take it away. On the quarterly deck it is still filed under Pilots. That gap is where AI adoption in market research stops.
Key Takeaways
- Roughly 60 percent of organizations evaluated a generative AI tool, 20 percent piloted and 5 percent reached production, per MIT NANDA, July 2025.
- That report names learning as the central barrier, ahead of infrastructure, regulation and talent: most deployed systems never retain feedback.
- Of 25 attributes McKinsey tested, workflow redesign mattered most for reaching operating profit. Only 21 percent had done any of it.
- A seat bought is not adoption. Adoption is one step of the research cycle done differently, with a named owner and a measurement.
- Staying in pilot is right when the method changes every project, when volume is low, or when the step feeds a regulated submission.
What counts as a pilot, and what counts as production
A pilot is a bounded trial of a tool on real work, with an end date, a small group of users, and no obligation to anyone outside that group. If it stops on a Tuesday, the study still ships.
Production is the state in which one step of the research cycle runs through the tool by default, on every study of that type, with a named owner and a number attached, such as days from field close to first readout. Stop it on a Tuesday and somebody notices before lunch.
The MIT NANDA report on the state of AI in business is open about its methodology: more than 300 public AI initiatives reviewed, 153 senior leaders surveyed. Roughly 95 percent were getting no measurable return on generative AI. A tool nobody depends on cannot show a return, because nothing changed that anyone measured.
How research teams run AI pilots today
This is what a competent team does when a promising tool lands.
- Someone senior signs up after a conference demo, on a departmental card.
- Two or three curious people use it on live client work, with no spare study to practise on.
- It works on one task: coding open ends, drafting a guide, updating a wave report.
- A short internal note gets written, and it becomes the only place the method exists.
- Legal asks where the data goes, and the answer arrives too late to stop the trial.
- The pilot settles into one person's routine, invisible on the process map.
Step three is not random. The GRIT Insights Practice Report 2026 records three tasks where agentic AI is already embedded in research: analyzing data, updating reports, and preparing and integrating data. Agentic AI means software that carries out a multi-step task on its own once given the goal. All three repeat every wave, which is where a pilot is easiest to graduate.
The numbers on pilots that never graduate
The funnel is public and it is unkind. In the MIT NANDA data, roughly 60 percent of organizations evaluated a generative AI tool, about 20 percent reached a pilot, and about 5 percent reached production. Nine out of ten evaluations never become something a colleague relies on.
Forbes, reporting the IDC CIO Playbook 2025, arrives from another direction with roughly 12 percent of pilots reaching production, and puts the loss down to organizational causes rather than technical ones.
PERSONAL EXPERIENCE The room matches the funnel. In our maturity conversations the tool is already bought and the pilot already ran. Two people use it well, nobody else has a reason to start, and nobody can name the number that would say whether it worked. The stall sits in the step after the pilot, the one nobody was assigned. Source: Cassi.ai.
The figure sets those stages side by side, 60 percent evaluated, 20 percent piloted, 5 percent in production, plus a fourth that is ours. The gate between the third and the fourth is what the rest of this article is about.
Where AI pilots stop, and the gate before the study calendar
A four stage funnel read from left to right. Evaluated holds about 60 percent of organizations, piloted about 20 percent, and in production about 5 percent, per the MIT NANDA State of AI in Business report of 2025. The fourth stage, embedded in the study calendar, carries no public number and is the criterion this article proposes. A gate stands between the third and fourth stage, and the gate is three questions about the step: who owns it by name on the process map, what number is measured on it and what that number was before, and what happens to it on the next wave when the person who ran the pilot is away.
The first three stages come from the MIT NANDA State of AI in Business report, 2025. The fourth stage and the gate are the criterion this article proposes.
Why the barrier is not the model
Teams read a stalled pilot as evidence that the model was not good enough, then spend a quarter evaluating a different one. The MIT NANDA report says that is the wrong place to look. The central barrier it identifies is learning: most deployed systems do not retain feedback, do not adapt to context and do not improve with use. Infrastructure, regulation and talent rank below that.
The consequence is concrete. The analyst corrects the same three code frame errors every wave, and the system produces them again on the next one. That is not a model quality problem. It is the absence of any mechanism carrying a correction forward, and no vendor sells that mechanism.
What actually correlates with impact
McKinsey tested 25 attributes of how organizations run AI to see which moves results. The winner was workflow redesign, meaning a change in the order and the ownership of the steps rather than in the software. In The state of AI: how organizations are rewiring to capture value, that attribute had the largest effect on whether AI showed up in EBIT, earnings before interest and taxes. Only 21 percent had fundamentally redesigned any of their workflows.
The same survey put oversight of AI governance with the chief executive as the element most correlated with reported impact. Nobody redesigns the order of the work by buying a license.
The unit of adoption is the step, not the seat
UNIQUE INSIGHT Most adoption dashboards count the wrong object. They count seats, logins, prompts sent, sometimes hours claimed as saved. None of that is adoption. Adoption is a step of the research cycle now performed differently, with a named owner and one number somebody actually looks at.
Take the eight stages of a study: brief, design, programming, field, processing, coding, analysis, report. Ask which is done differently now than two years ago. The honest answer is usually one, sometimes none. The stages that graduate first are [questionnaire testing before field](/blog/survey-testing-before-fieldwork) and [the first pass over qualitative material](/blog/qualitative-coding-evidence-chain), because both repeat and both have somebody's name on them.
A pilot not fastened to a step has nothing to graduate into, which is why counting pilots is a losing measurement.
The two journeys, individual and organizational
UNIQUE INSIGHT Two things have to change at once, and most teams fund only the first.
The individual journey is what one researcher learns:
- Specify the task: what to ask the system for, and what to withhold.
- Learn where the output is reliably wrong, and check that part first.
- Record a correction as an edit to the instruction, not a fix to one file.
The organizational journey is what changes around that person:
- Name an owner for the step, not for the tool.
- Define the measurement, and write down the before number while it exists.
- Put the instruction somewhere shared and versioned a new hire can open.
- Give the step a slot in the study calendar, so it is scheduled rather than improvised.
- Decide who runs it when the owner is on leave.
Fund only the first list and the capability leaves with the person. That is the mechanism behind every pilot that quietly stopped working.
Where Cassi.ai comes in
Cassi.ai maps the two journeys that decide whether a pilot graduates, the individual one, where a researcher learns to do a specific step differently, and the organizational one, where that step gets an owner, a measurement and a place in the study calendar. That mapping is the AI Maturity Framework.
Cassi.ai is a software engineering company specialized in the pains of market research, innovation and insights, built by people with twenty years in the field. The systems behind it are in the Cassi.ai portfolio.
A pilot compared with a graduated step
| Criterion | Pilot | Graduated step |
|---|---|---|
| Who owns it | The person who set it up | A named role, written on the process map |
| What is measured | Enthusiasm, sometimes hours claimed as saved | One number, with a before value on record |
| When that person leaves | Reverts to manual within one wave | Someone else runs it from the written instruction |
| Where the instruction lives | A chat history and one person's memory | A shared, versioned document a new hire can open |
| On the next wave | Rebuilt from memory, slightly differently | Runs by default, and corrections carry forward |
Comparison built on the failure modes in the MIT NANDA report.
The governance gap that only shows up at scale
AI governance is the set of rules deciding who may use which system, on which material, with what record kept, and who answers when the output is wrong. Skipping it costs nothing while two people experiment. It becomes the binding constraint the moment one step runs on client material across every study.
The GRIT Insights Practice Report 2026 names AI governance as a critical gap in the research industry, and records a correlation between confidence in AI risk management and beating business targets. The material is confidential, the respondent data is regulated, and the question of which model saw which transcript has a right answer. That is why a general chat subscription does not become a research platform by being used heavily, covered in [the six gaps between a chatbot and an enterprise platform](/blog/why-chatgpt-is-not-a-research-platform), and why [one model alone leaves a blind spot](/blog/single-model-bias-market-research).
When staying in pilot is the right call
PERSONAL EXPERIENCE Not every step should graduate, and pretending otherwise is how a maturity framework turns into a sales document.
Three cases where we tell teams to leave it alone. A method that changes shape every project, where specifying it costs more than doing the work. Low volume, where a step running twice a year never repays an owner, a measurement and a calendar slot. And a regulated submission, where the audit requirement is a documented human decision.
The tell is the next wave. With none, there is nothing to graduate into, and a permanent pilot is the honest answer rather than the cowardly one.
FAQ
Why do most AI pilots fail to reach production?
Because the pilot was attached to a tool instead of a step. MIT NANDA names learning as the central barrier: systems that do not retain feedback repeat the same errors every wave, so no correction accumulates.
What does AI maturity mean for an insights team?
UNIQUE INSIGHT It means counting stages of the research cycle now done differently, each with an owner and a measurement, rather than counting tools or seats. One graduated step beats six live pilots.
How do you measure whether an AI pilot worked?
Pick one number on one step and record the before value while it exists: days from field close to first readout, or hours to a first code frame. McKinsey found workflow redesign the strongest of 25 attributes tested, so the measure belongs on the step.
Should a research team build its own AI tools or buy them?
MIT NANDA found tools bought from specialised vendors succeeded at roughly twice the rate of systems built internally. That is inconvenient for teams who value autonomy. Either way it matters less than whether the step has a named owner.
Does AI adoption in market research require redesigning the workflow?
Yes, in the sense that the order and the ownership of the steps has to change. McKinsey put workflow redesign at the top of 25 attributes tested, with only 21 percent having done any of it. A tool dropped into an unchanged sequence buys a faster version of it.
A pilot that has run for fourteen months has stopped being a pilot. It is a production system nobody has agreed to own. If your team still calls a tool a pilot after two study waves, name the step, name its owner, and write down the number that would tell you it worked.
Published by Cassi.ai. Read the full article at https://www.cassiai.com/blog/ai-adoption-in-market-research-past-the-pilot.