Skip to main content

Case study

In 1 of every 8 turns, this Zoom transcript changed what the participant said.

Zoom captured 96% of the words. 57% of speaking turns had no mistake. In 13% of turns, the participant's meaning changed.

A public health research team recorded an 86-minute focus group on Zoom and had its automatic transcript in hand within minutes. We compared that file, word for word, against the analysis-ready transcript our team delivered for the same recording. This is what we found.

n = 1 focus group · 86 min · 12 participants · 9,512 words · 141 turns · 740 differences, each classified

96%
of words captured
9,125 of 9,512 words in the final transcript appear in Zoom’s file
43%
of speaking turns contain at least one mistake
60 of 141 turns
13%
of turns contain a mistake that changes the meaning
19 of 141 turns
24
mistakes that change what a participant said
of 109 real transcription mistakes

The setup

If you've run a focus group on Zoom recently, you've seen the transcript it produces. It arrives fast, it's free, and it comes with speaker names and timestamps already attached. For a lot of teams the question isn't whether to use it, but how much cleanup it needs before it can go into NVivo or Dedoose.

We had an unusual chance to answer that with real numbers. A team of public health professionals held a 12-person focus group about a proposed clinical program in their field. The moderators had Zoom's transcript. We had the audio, and we produced a clean-verbatim, de-identified transcript the way we would for any qualitative project. Then we lined the two up.

What Zoom did well

We'll start here because it matters, and because the free transcript is better than its reputation in some ways.

  • It captured 96 of every 100 words. Of the 9,512 words in the final transcript, 9,125 appear in Zoom's file, in order. Zoom hears English well.
  • It labeled the speaker correctly on every line we could check. 137 turns, 137 right. Zoom tags speech by whoever is logged in, so when each person is on their own device it doesn't get confused.
  • It kept 10 participants separate. The study only needed respondents tracked as a group, so Zoom recorded more speaker detail than the project asked for.

What it got wrong, and where

Zoom made 109 real transcription mistakes over 86 minutes, roughly one every 47 seconds. That number on its own isn't alarming. What matters is where they landed.

Since a turn is what a researcher actually quotes, the turn-level numbers are the ones to think about. Pull a passage from this transcript at random and there's a 4-in-10 chance something in it is wrong, and about a 1-in-8 chance it's wrong in a way that would change how you code it.

The 24 meaning-changing errors weren't scattered across ordinary conversation. They clustered on exactly the vocabulary the study is about.

Zoom wroteWhat was saidWhat it is
the analysis (3 times)dialysisThe field the whole conversation was about
ferret in / cortisoneferritin / cortisolBiomarkers
jean impressiongene expressionCentral concept
farm assists / social walkerspharmacists / social workersWho is being discussed
insureduninsuredOpposite sense
three300Off by a factor of 100
we screen thembe screened inChanges who is acting
weighed / assistwait / insistWrong verb
a spire (twice)ASPIREThe study’s own acronym
sarcoid oasissarcoidosisOne of the conditions under discussion

The words in this table are same-class substitutes for the originals; see the confidentiality note at the end.

Every one of the 24 is a real English word in a grammatical sentence. Spell-check won't flag them. Reading alone can miss them. Finding them reliably means listening.

That's the part that makes a Zoom transcript different from a transcript with typos. A typo announces itself. “Ferret in” in a sentence about iron studies does not, unless you already know the answer, and the person cleaning the transcript often isn't the person who does.

What the raw file looks like

Beyond the words, Zoom hands you 628 timestamped fragments, each a few seconds long. The finished transcript has 141 speaker turns. Somebody has to stitch the fragments back into a conversation before anyone can read it, let alone code it.

Two speaker problems also needed a person to sort out. One participant with 27 turns was labeled with the name of their phone, since that's what they'd logged in as. And for one stretch, two people spoke through a single microphone and Zoom attributed all of it to one name. Neither is a transcription error exactly, but both would have gone into the analysis wrong.

What it took to fix

We counted every difference between Zoom's file and the finished transcript. There were 740 of them, about 8.6 per minute of audio.

Type of editCountShare
Fillers, stutters and false starts removed (clean verbatim)44460%
Names and locations replaced with <Name> and <State>10014%
Spelling conventions (going to → gonna)8011%
Zoom’s transcription mistakes corrected10915%
Other71%

So most of the edits, about 85%, are making the file read like a research transcript. That part is tedious but safe; you can see a filler word. The other 15% is the part you can't see, and it's the reason the file has to be checked against audio rather than skimmed. Edit counts are not effort: the time estimate below weights them.

What that means in hours

We built up an estimate for a research coordinator doing this carefully with Word and a media player, based on the edits above.

StepMinutes
Listen to the full recording once86
Stop, rewind and correct 109 mistakes64
Remove 444 fillers and false starts22
De-identify 100 names and locations8
Rebuild 628 fragments into turns, resolve speakers25
Final read and formatting19
Total~225, about 3¾ hours

Roughly 2.6 times the length of the recording. This is an estimate built from the edit counts, not a stopwatch.

The common advice online is to budget 15 to 30 minutes of review per hour of audio for an AI transcript. On this file, that pace would have caught the fillers and likely left most of the 24 meaning-changing errors in place, because catching them usually means hearing them.

A note on our own work. When we ran this comparison we also found 6 lines in our delivered transcript where a moderator's question had been labeled as a respondent. Zoom had those right. We've corrected them and mentioned it here because the point of the exercise was to be accurate about both files, not just one.

When the Zoom transcript is enough

It'd be dishonest to suggest every recording needs this treatment. Zoom's transcript is probably fine when

  • you need to find a moment in the recording, not quote it,
  • the conversation is everyday language with little specialist vocabulary,
  • there are two speakers on separate devices, and
  • nothing in it is going to be coded, quoted, or shared beyond the team.

This focus group was the opposite on every count. Twelve people, clinical vocabulary throughout, participants' names and their location spoken aloud, and a transcript that would be coded and cited. Given that, the 4% Zoom missed was the 4% the study most needed.

What we delivered

The same recording, transcribed by a person against the audio, in clean verbatim, with every name and location replaced by a tag, moderator and respondents labeled, and pauses and crosstalk marked. It's the file that goes into the analysis software without anyone on the research team having to listen to the recording again.

How we measured. Zoom's .vtt file and our delivered .docx were compared word by word after removing punctuation and case. Each of the 740 differences was classified as a style edit, a de-identification, or a transcription mistake, and mistakes were rated on whether a reader of the Zoom version would take away something the participant didn't say. Word and turn counts are exact. The 96% is word matching: the share of words in the final transcript that also appear, in order, in Zoom's file. It is not a conventional word error rate, which would also charge Zoom for substitutions, insertions and deletions and would come out lower. The reference for “what was said” is Landmark's delivered transcript, produced by a transcriptionist working from the recording and proofread by a second person. For this comparison, a proofer listened to the recording for each of the 24 meaning-changing items; the remaining differences were classified from the two texts, which is why 109 carries a range. The 6 speaker-label errors in our own file were found the same way and corrected. A different reviewer might move a handful of borderline items, so read 109 as roughly 100 to 115 and 24 as roughly 20 to 28. The coordinator time is an estimate built from the edit counts. This is one file, and a hard one: twelve speakers and clinical vocabulary. Zoom will do better on a two-person interview in everyday language.

Confidentiality. Study identity, field, participant names and location have been withheld, and the specific words in the examples have been changed. Each substitute is the same kind of word as the original (a biomarker for a biomarker, a job title for a job title, a study acronym for a study acronym) and produces the same kind of error, so the examples show exactly what happened without allowing the study to be identified.