Case study
In 1 of every 8 turns, this Zoom transcript changed what the participant said.
Zoom captured 96% of the words. 57% of speaking turns had no mistake. In 13% of turns, the participant's meaning changed.
A public health research team recorded an 86-minute focus group on Zoom and had its automatic transcript in hand within minutes. We compared that file, word for word, against the analysis-ready transcript our team delivered for the same recording. This is what we found.
n = 1 focus group · 86 min · 12 participants · 9,512 words · 141 turns · 740 differences, each classified
The setup
If you've run a focus group on Zoom recently, you've seen the transcript it produces. It arrives fast, it's free, and it comes with speaker names and timestamps already attached. For a lot of teams the question isn't whether to use it, but how much cleanup it needs before it can go into NVivo or Dedoose.
We had an unusual chance to answer that with real numbers. A team of public health professionals held a 12-person focus group about a proposed clinical program in their field. The moderators had Zoom's transcript. We had the audio, and we produced a clean-verbatim, de-identified transcript the way we would for any qualitative project. Then we lined the two up.
What Zoom did well
We'll start here because it matters, and because the free transcript is better than its reputation in some ways.
- It captured 96 of every 100 words. Of the 9,512 words in the final transcript, 9,125 appear in Zoom's file, in order. Zoom hears English well.
- It labeled the speaker correctly on every line we could check. 137 turns, 137 right. Zoom tags speech by whoever is logged in, so when each person is on their own device it doesn't get confused.
- It kept 10 participants separate. The study only needed respondents tracked as a group, so Zoom recorded more speaker detail than the project asked for.
What it got wrong, and where
Zoom made 109 real transcription mistakes over 86 minutes, roughly one every 47 seconds. That number on its own isn't alarming. What matters is where they landed.
Since a turn is what a researcher actually quotes, the turn-level numbers are the ones to think about. Pull a passage from this transcript at random and there's a 4-in-10 chance something in it is wrong, and about a 1-in-8 chance it's wrong in a way that would change how you code it.
The 24 meaning-changing errors weren't scattered across ordinary conversation. They clustered on exactly the vocabulary the study is about.
| Zoom wrote | What was said | What it is |
|---|---|---|
| the analysis (3 times) | dialysis | The field the whole conversation was about |
| ferret in / cortisone | ferritin / cortisol | Biomarkers |
| jean impression | gene expression | Central concept |
| farm assists / social walkers | pharmacists / social workers | Who is being discussed |
| insured | uninsured | Opposite sense |
| three | 300 | Off by a factor of 100 |
| we screen them | be screened in | Changes who is acting |
| weighed / assist | wait / insist | Wrong verb |
| a spire (twice) | ASPIRE | The study’s own acronym |
| sarcoid oasis | sarcoidosis | One of the conditions under discussion |
The words in this table are same-class substitutes for the originals; see the confidentiality note at the end.
Every one of the 24 is a real English word in a grammatical sentence. Spell-check won't flag them. Reading alone can miss them. Finding them reliably means listening.
That's the part that makes a Zoom transcript different from a transcript with typos. A typo announces itself. “Ferret in” in a sentence about iron studies does not, unless you already know the answer, and the person cleaning the transcript often isn't the person who does.
What the raw file looks like
Beyond the words, Zoom hands you 628 timestamped fragments, each a few seconds long. The finished transcript has 141 speaker turns. Somebody has to stitch the fragments back into a conversation before anyone can read it, let alone code it.
Two speaker problems also needed a person to sort out. One participant with 27 turns was labeled with the name of their phone, since that's what they'd logged in as. And for one stretch, two people spoke through a single microphone and Zoom attributed all of it to one name. Neither is a transcription error exactly, but both would have gone into the analysis wrong.
What it took to fix
We counted every difference between Zoom's file and the finished transcript. There were 740 of them, about 8.6 per minute of audio.
| Type of edit | Count | Share |
|---|---|---|
| Fillers, stutters and false starts removed (clean verbatim) | 444 | 60% |
| Names and locations replaced with <Name> and <State> | 100 | 14% |
| Spelling conventions (going to → gonna) | 80 | 11% |
| Zoom’s transcription mistakes corrected | 109 | 15% |
| Other | 7 | 1% |
So most of the edits, about 85%, are making the file read like a research transcript. That part is tedious but safe; you can see a filler word. The other 15% is the part you can't see, and it's the reason the file has to be checked against audio rather than skimmed. Edit counts are not effort: the time estimate below weights them.
What that means in hours
We built up an estimate for a research coordinator doing this carefully with Word and a media player, based on the edits above.
| Step | Minutes |
|---|---|
| Listen to the full recording once | 86 |
| Stop, rewind and correct 109 mistakes | 64 |
| Remove 444 fillers and false starts | 22 |
| De-identify 100 names and locations | 8 |
| Rebuild 628 fragments into turns, resolve speakers | 25 |
| Final read and formatting | 19 |
| Total | ~225, about 3¾ hours |
Roughly 2.6 times the length of the recording. This is an estimate built from the edit counts, not a stopwatch.
The common advice online is to budget 15 to 30 minutes of review per hour of audio for an AI transcript. On this file, that pace would have caught the fillers and likely left most of the 24 meaning-changing errors in place, because catching them usually means hearing them.
When the Zoom transcript is enough
It'd be dishonest to suggest every recording needs this treatment. Zoom's transcript is probably fine when
- you need to find a moment in the recording, not quote it,
- the conversation is everyday language with little specialist vocabulary,
- there are two speakers on separate devices, and
- nothing in it is going to be coded, quoted, or shared beyond the team.
This focus group was the opposite on every count. Twelve people, clinical vocabulary throughout, participants' names and their location spoken aloud, and a transcript that would be coded and cited. Given that, the 4% Zoom missed was the 4% the study most needed.
What we delivered
The same recording, transcribed by a person against the audio, in clean verbatim, with every name and location replaced by a tag, moderator and respondents labeled, and pauses and crosstalk marked. It's the file that goes into the analysis software without anyone on the research team having to listen to the recording again.
How we measured. Zoom's .vtt file and our delivered .docx were compared word by word after removing punctuation and case. Each of the 740 differences was classified as a style edit, a de-identification, or a transcription mistake, and mistakes were rated on whether a reader of the Zoom version would take away something the participant didn't say. Word and turn counts are exact. The 96% is word matching: the share of words in the final transcript that also appear, in order, in Zoom's file. It is not a conventional word error rate, which would also charge Zoom for substitutions, insertions and deletions and would come out lower. The reference for “what was said” is Landmark's delivered transcript, produced by a transcriptionist working from the recording and proofread by a second person. For this comparison, a proofer listened to the recording for each of the 24 meaning-changing items; the remaining differences were classified from the two texts, which is why 109 carries a range. The 6 speaker-label errors in our own file were found the same way and corrected. A different reviewer might move a handful of borderline items, so read 109 as roughly 100 to 115 and 24 as roughly 20 to 28. The coordinator time is an estimate built from the edit counts. This is one file, and a hard one: twelve speakers and clinical vocabulary. Zoom will do better on a two-person interview in everyday language.
Confidentiality. Study identity, field, participant names and location have been withheld, and the specific words in the examples have been changed. Each substitute is the same kind of word as the original (a biomarker for a biomarker, a job title for a job title, a study acronym for a study acronym) and produces the same kind of error, so the examples show exactly what happened without allowing the study to be identified.