Someone is going to pitch you automated qualitative analysis this year. Maybe it’s a platform. Maybe it’s a research firm running one quietly behind the scenes. The pitch will be faster turnaround at lower cost. The deliverable will look like every findings deck you’ve ever approved.
Here’s the problem: you have almost no way to tell from the deck whether the analysis actually happened.
A PLOS ONE study tested this directly. Researchers handed Microsoft Copilot a set of interview transcripts and asked for a thematic analysis: themes, participant counts, supporting quotes, all inside a tight word limit. Copilot delivered exactly that. Confident, well-formatted, ready to drop into a slide.
Most of it was built from the first two or three pages of the data. Some of the quotes were never in the transcripts at all.
Every component a findings section is supposed to have, wrapped around analysis that may never have happened. We call that the rigor costume.
Why “faster coding” is solving the wrong problem
A companion paper, from Kien Nguyen-Trung and Susanne Friese in the International Journal of Social Research Methodology, gets at the mechanics. It also hands you language for a conversation with a vendor that doesn’t require anyone to argue about whether AI is good or bad. The real question is narrower than that, and more answerable.
Start here: why does qualitative research involve coding at all? No one can hold twenty hours of interviews in working memory, so we break transcripts into pieces, label them, and use those labels to find patterns later. That practice is close to a century old. It’s a workaround for a limit human researchers have.
A language model doesn’t have that same limit. Within the material it can actually see, it can compare passages directly, without converting each one into a tag first.
Which means a tool built to mimic manual coding, faster, is solving a problem the tool doesn’t have. That matters commercially, because coding is the easiest thing to sell. It’s the visible, billable, tedious part of the job. Speeding up the tedious part isn’t the same as making the analysis better. A vendor whose whole pitch is coding speed is selling you the wrong improvement.
Not all qualitative research is the same kind of rigorous
This is the part worth pulling out of academic language, because it determines what you should expect from any analysis, human or machine.
Qualitative research runs along a range. Methodologists have sorted that range into three broad zones for decades. Which zone a study sits in changes what counts as a finding, what counts as rigor, and what a machine can reasonably contribute.
Picture one subject explored three ways: enterprise adoption of an AI coding assistant.
Small q. A developer survey. Six hundred written answers to “what slowed your team’s adoption?” sorted into categories defined in advance, reported as percentages. “41% cite security review delays” on a slide. Consistency between two coders is a meaningful quality check here, because the whole premise is that answers can be classified the same way twice.
Medium Q. Interviews with engineering managers at accounts where seat growth stalled. A discussion guide built from client hypotheses plus judgment about what else is worth probing. Structure going in, interpretation inside it. An analyst notices that “security concerns” at a regulated bank and “security concerns” at a fast-moving startup aren’t the same finding wearing the same two words. One is a real compliance gate. The other is an objection nobody wants to argue with.
Big Q. Long, open conversations with senior engineers about what these tools are doing to their sense of craft. No rigid framework going in. The finding is a pattern nobody said out loud, surfaced by an analyst reading for what sits underneath what was said. You don’t get there from a code count. You get there from an analyst with enough context to hear it — and it’s the kind of finding that changes how a client runs the rollout.
Most commercial B2B research lives in medium Q, and that’s a legitimate methodological choice, not a compromise. But the highest-value findings in AI adoption work often sit at one end or the other: a clean number worth defending in a room, or an interpretation nobody said out loud because organizations are political and people protect themselves.
Here’s why the zones matter when you’re evaluating a vendor. The most common failure is mixing them up: a study claims deep interpretive work, then validates it by measuring how often the model and a human agreed on a label. That check belongs to small q, where consistency of classification is the point. Bolted onto a method whose whole premise is that interpretation is situated and subjective, what you get is borrowed rigor — a validity check answering a question the study never asked.
Five ways this goes wrong
Invented quotes. The one that ends client relationships. Ask a model for themes, then for quotes to support them. If the source passage isn’t actually retrievable at the moment the answer gets generated, the system can produce something that sounds exactly like a real respondent. Nobody catches it in review, because it’s precisely what a real person would say. Every quote in a deck should trace back to a specific line in a transcript.
Front-of-file bias. In the Copilot study, most themes and quotes came from the first few pages of a much longer document. Your 30-interview study can quietly start behaving like a five-interview study, with no signal that it happened.
Topics dressed up as themes. A model hands you “Security Concerns” and “Integration Complexity.” Those are buckets, not findings. A theme carries an argument inside it — something closer to “security review has become the acceptable way for reluctant managers to veto a tool they don’t want for reasons they won’t say out loud.” Buckets fill a slide. A real theme changes what a client does next.
Merging on labels instead of meaning. Ask a model to tidy up a code list and it compares names and short descriptions. “Procurement friction” and “buying committee drag” might be the same idea or two different ones. A tidy label like “governance” can quietly absorb the distinction that was carrying the whole finding. Someone has to go back to the actual passages to decide what those respondents meant.
It averages. Pattern-seeking systems surface what repeats more easily than the one account that changes the interpretation. In AI adoption research, the value is sometimes in the single respondent who’s eighteen months ahead of everyone else. A conventional summary can make that person look like noise. A good analyst asks whether he’s actually the leading edge.
Where AI genuinely earns its place
Pointed at the right work, the value is real:
- Finding things, with receipts. Every mention of security review across 28 transcripts, tied back to source.
- Attacking your own theme. An analyst brings a tentative read, then uses the model to hunt for passages that complicate or contradict it.
- Comparing segments. Where regulated-industry respondents diverge from software-native ones, or where a respondent contradicts herself twenty minutes later.
- Finding the holes in the guide. What did we fail to ask, given what participants kept raising unprompted?
- Forcing grounding. Working one transcript at a time, then reconciling across interviews, is a defensible way to stop a model from pretending it engaged with material it never actually processed.
The highest-value use of these is attacking your own theme. Most AI workflows are built to generate answers. Qualitative analysis usually gets better when you use the system to generate problems for your answer instead.
Six questions for your next vendor conversation
Useful whether you’re evaluating a research firm or a platform:
- What kind of study are you claiming, and does your quality check match it? If a vendor describes deep interpretive analysis and validates it mainly through model-human agreement, ask how those two things fit together.
- Can every quote in the deck trace back to a line in a transcript? Ask to see it done live, on one quote.
- How much data did the model actually have access to, and how do you know? A context-window number isn’t the answer. Ask how the system retrieves source material and how it verifies coverage.
- Who cleaned up the code list, and did they go back to the passages to do it? Two similarly named codes may be doing different analytic work.
- What does the audit trail look like? Prompts kept. Outputs accepted and rejected. Decisions recorded. Source passages preserved, dated. A chat log by itself isn’t an audit trail.
- When was that capability tested, and on what system? This cuts both ways. A vendor citing a benchmark from two model generations back may be telling you very little. And a skeptic treating last year’s chatbot behavior as a permanent property of every model is making the same mistake in reverse.
The judgment doesn’t come from the tool
None of the raw material an analyst works with settles the interpretation on its own. Not the transcripts, not a framework, not a colleague’s read, and not a model’s suggestions. Someone still has to build the finding.
That’s always been the actual deliverable: the judgment to notice which apparent outlier changes how you read the other 29, the discipline to tell a real pattern from one you’re forcing, and walking away from a great-sounding theme the data doesn’t back up. That’s the actual deliverable, not the interview count or the coding hours.
This is exactly the kind of judgment we mean when we talk about Fluid Intelligence®® — the ability to aim AI at the right part of a problem, give it the right context, and challenge it when it’s wrong. It’s not a soft skill bolted onto a research process. It’s the difference between a finding and a rigor costume.
If you’re not sure whether your last AI-assisted study has this problem, that’s worth a conversation. We build AI into our own qualitative work every day, in the places where it holds up, with a human accountable for every finding that reaches a client. See how we approach our B2B market research studies.
If your team is the one running these tools in-house and you want a way to tell real analysis from a rigor costume before it reaches a client or an exec, that’s what we teach in AI Training for Teams and build into every AI Strategy Sprint. Let’s Talk.