Qualitative Coding With AI Without Losing the Evidence
Qualitative coding with AI shortens the first pass over interview transcripts. How open coding runs today, what it costs, and where review time returns.
A researcher opens the transcript of interview number seven. She reads paragraph by paragraph, marks the passage where the respondent hesitates before saying the price felt fair, writes a label next to it, and moves on. Two hundred pages later she has four hundred labels, and the real work starts, which is deciding which of those labels are the same thing wearing different words.
Key Takeaways
- Open coding of interviews is still manual in most research teams, including teams that already use language models for other tasks.
- The bottleneck is the second pass, where hundreds of spreadsheet rows get merged into themes by hand.
- The cost lands on the deadline rather than on the hours sheet, and the insight lead is the one who explains it to the client.
- An assisted first pass only pays off above roughly 40 interviews. Below that, review time eats the gain.
- Irony and one-word answers are where machine coding fails hardest, and where checking costs more than coding from scratch.
What open coding and thematic analysis actually mean
Open coding is the first pass over qualitative material, where the researcher reads without a predefined category list and attaches a short descriptive label to every passage that carries meaning. The labels come from the data instead of from a prior framework.
Thematic analysis is the method that turns those labels into findings. It identifies, reviews and names patterns across the whole data set rather than reporting each interview separately. The six-phase framework most teams follow runs from familiarization with the data, through generation of codes and combining codes into themes, to reviewing themes and reporting findings, as set out in Braun and Clarke's framework summarized by National University.
The vocabulary matters because the two steps have different failure modes. Coding fails by inconsistency. Thematic analysis fails by losing the trail back to the passage that started the theme.
How coding is done today, step by step
Nothing here is exotic. It is what a competent team does on a Tuesday.
- Export the transcripts, one document per session, usually after a transcription service returns them.
- Read each transcript in full once, without coding, to get the shape of the conversation.
- Read it again, marking passages and writing a label beside each one, in the document margin or in a coding tool.
- Dump every label into a spreadsheet, one row per coded passage, with the participant, the session, and the quote.
- Sort the spreadsheet and merge duplicate labels by hand, because "price is high" and "too expensive for what it is" arrived as separate rows.
- Group the surviving labels into candidate themes, then argue about the grouping with a second researcher.
- Go back to the transcripts to find the quote that will carry each theme in the report.
Step five is where the study lives or dies. A 30 interview study produces a spreadsheet with several hundred rows, and someone reviews every one of them by hand. When a second coder is involved, the team also checks how far the two of them agree. Intercoder reliability is a measure of how consistently two coders apply the same label to the same passage, and it is calculated with statistics such as Cohen's kappa, which adjusts for agreement that would happen by chance. It has known limits: it works with two coders only, and it is documented as overly conservative in some designs.
What it costs, and who answers for it
UNIQUE INSIGHT The cost of manual coding almost never shows up as a budget problem, because the hours were already inside the fee. It shows up as a schedule problem. The transcripts land late, the coding takes the week that was reserved for interpretation, and the readout gets built in two days by people who are tired.
The client does not see the spreadsheet. The client sees a debrief that arrives eleven days after fieldwork closed, and a recommendation that sounds thinner than the fieldwork felt. The insight lead is the person in the room for that conversation, and nobody in that room is discussing coding methodology. They are discussing whether the deadline is realistic next wave.
That is the pain worth naming. The interpretation phase, which is the part clients are actually paying for, gets compressed into whatever time the mechanical phase leaves behind.
What an evidence chain is, and why it breaks
An evidence chain is the traceable path from a raw passage in a transcript to the recommendation that ends up on a slide. It has four links: evidence, interpretation, implication, recommendation. The chain is intact when you can start at the recommendation and walk backwards to the exact thing a participant said.
UNIQUE INSIGHT Most qualitative workflows break the chain at the merge step. Once "price is high" and "too expensive for what it is" collapse into one theme label, the pointer back to the source disappears. Who said it, in which session, under which stimulus, now lives only in the memory of the person who did the merging. Six weeks later, when the client asks how many participants actually raised cost, the answer is a reconstruction rather than a lookup. This is also why [summarizing a transcript with a general purpose chatbot](/blog/why-chatgpt-is-not-a-research-platform) solves the wrong half of the problem. It compresses the material fast and discards provenance at the same speed.
Where Cassi.ai comes in
PERSONAL EXPERIENCE Qual Lab is the system we built for this. The material goes in as transcripts, documents, images, audio or video, gets split into retrievable chunks, and then a sequential orchestrator runs 19 analysis modules over it. Each module is a different analytical lens with its own retrieval query rather than one generic summary prompt: synthesis, themes, emotion, language, drivers, tensions, journey, clusters, semiotics, contradictions, temporality, stimulus response, evidence matrix and dashboard narrative, among others.
PERSONAL EXPERIENCE Each module carries a confidence estimate built from three inputs: how much evidence it found, how much of the participant base that evidence covers, and how rich the resulting output is. A theme drawn from two people out of twenty says so on its face. Every insight record keeps its link to the underlying passages, plus its own limitations field, so the four links of the chain stay addressable after the merge.
Cassi.ai is a software engineering company specialized in the pains of market research, innovation and insights, built by people with twenty years in the field. Qual Lab is the piece of that work aimed at qualitative analysis, alongside the same approach applied to [questionnaire testing before fieldwork](/blog/survey-testing-before-fieldwork) and to [running an analysis across several models instead of one](/blog/single-model-bias-market-research).
Manual coding compared with an assisted first pass
| Criterion | Fully manual coding | Assisted first pass with human review |
|---|---|---|
| Where the effort sits | Reading, labeling and merging | Reviewing and correcting labels |
| Second pass on a 30 interview study | Several hundred spreadsheet rows merged by hand | Candidate themes arrive grouped, with source passages attached |
| Traceability after merging | Depends on the coder's notes and memory | Held in the record, queryable later |
| Handling of irony and short answers | Reliable, because a human hears tone | Weakest point, needs line by line checking |
| Consistency across a long study | Drifts as the coder tires | Stable, and wrong in the same way throughout |
PERSONAL EXPERIENCE Comparison drawn from how the Qual Lab pipeline behaves against the manual workflow it replaces. Source: Cassi.ai.
The last row deserves attention. Machine coding is consistent, which sounds like an advantage and is only half of one. A tired human coder drifts in random directions, so the errors partly cancel. A model that misreads a construction misreads it identically in all thirty transcripts, and a systematic error is harder to spot in review than a scattered one.
Where this gets it wrong
PERSONAL EXPERIENCE Two failure cases are worth stating plainly, because they decide whether the approach is worth running at all.
Irony is the first. A respondent who says "oh, fantastic, another app that wants my location" is coded as positive sentiment by a model reading the words. A human coder hears the eye roll. Fixing that kind of label requires going back to the audio, and the reviewer ends up doing the original coding work plus the work of undoing a wrong label.
Short answers are the second. "Yeah." "I guess." "Sometimes." These carry meaning that lives entirely in what preceded them and in how they were delivered. Machine labels on that material are guesses dressed as codes, and checking them is slower than coding them from scratch.
The consequence is a threshold. Below roughly 40 interviews, review time cancels the gain. The benefit appears in volume, in recurring waves, and in studies where somebody will ask six months later which participants said what.
When to automate the first pass and when not to
| If your situation is | Then | Because |
|---|---|---|
| Under 40 interviews, one-off study | Keep coding manually | Setup and review time cancel the gain |
| Above 40 interviews, recurring waves | Automate the first pass, review by hand | The gain compounds across waves |
| Heavy irony, very short answers, vox pop | Human coding | The failure mode is the material itself |
| Regulated or legal submission | Human coding, documented | Auditability outranks speed |
| Multi-format material: audio, video, stimulus boards | Assisted, with the evidence chain preserved | Manual traceability across formats is where teams lose the trail |
UNIQUE INSIGHT The deciding variable is not study size on its own. It is whether anyone will need to reopen the study later. A one-off tracker readout can survive a broken evidence chain. A brand study that feeds three years of decisions cannot.
FAQ
Does qualitative coding with AI replace the researcher?
No. It replaces the first pass over the transcript and the mechanical part of the merge. The interpretation, the argument about what the themes mean, and the recommendation stay with the researcher. UNIQUE INSIGHT The realistic gain is that the interpretation phase stops being the part that gets compressed when the schedule slips.
How many interviews justify automating the first pass?
Roughly 40 in our experience, and the number moves with the material. PERSONAL EXPERIENCE Below that, the review of machine labels costs about what the coding would have cost. Above that, and especially across recurring waves, the setup amortizes.
Can I still check intercoder reliability if a model does the first pass?
Yes, by treating the model as one coder and a human as the other, then computing agreement as usual. The caveat is that Cohen's kappa works with two coders only and can be conservative, so read the number as a signal about label definitions rather than as a verdict.
What exactly gets lost when the evidence chain breaks?
The ability to answer "who said this, and where". Once labels are merged without preserved pointers, the link between a recommendation and the passage behind it becomes a reconstruction from memory. That is the part that fails under client scrutiny six weeks after the debrief.
Is a general purpose chatbot enough for this?
It is enough to summarize one transcript. It does not hold a participant base, it does not track how much of that base supports a theme, and it does not keep a passage-level record you can query later.
Qualitative coding is the phase where research teams still spend their most senior hours doing their least senior work, and the client only ever sees the consequence of that on the calendar. If your team still merges coding spreadsheets by hand, the number worth measuring this quarter is how many days of interpretation time that merge consumed on your last study of 40 interviews or more.
Published by Cassi.ai. Read the full article at https://www.cassiai.com/blog/qualitative-coding-evidence-chain.