Skip to main content

Zoom, Teams, Otter, Dragon & AI transcripts

Zoom or Otter transcript not accurate enough? Send it with the recording. We correct it.

Research teams tell us the same thing every week: the auto-generated transcript is roughly right, they are spending hours per audio hour correcting it, and they still don't trust the result. That correction is work we do. Send the transcript and the recording, and a proofreader checks it word by word against the audio: wrong words, missed speech, speaker labels, sentence boundaries, and the names and terminology the tool guessed.

AI vendors are careful with their wording: the promise is “up to 99% accuracy on clear audio” (Sonix pricing FAQ, August 2026). A caregiver interview recorded in a hospital cafeteria, a focus group talking over itself, a participant with a soft voice or an accent the model wasn't trained on — that is not clear audio, and it is what research data usually sounds like. Our 99%+ standard for transcription applies to the recordings researchers actually send us, not to a studio demo.

Two ways to order. Review and correction of your existing transcript is quoted per project once we have seen the transcript and the recording. Transcription from the recording is the standard rate: from $1.99 per audio minute for one-on-one interviews and from $2.69 per audio minute for focus groups and recordings with three or more speakers, whether or not a draft exists. A mostly-right draft takes a proofreader less time than one that has to be retyped; if starting from the audio would cost the same or less, the quote says so.

What we correct

A transcript you can code, quote, and put in a methods section has to be checked against the audio word by word. The proofreader listens to the entire recording with your draft open, and fixes the problems that turn up in nearly every auto-generated transcript of a research recording:

Speaker attribution
Drafts merge turns, split one speaker into several, or label the wrong person; in focus groups the labels are rarely usable. Every turn is re-listened to and re-assigned.
Overlapping talk
When two people speak at once, the draft keeps one of them or neither. A research transcript marks [crosstalk] and captures each voice where it is audible.
Names, places, terminology
Participant names, medications, instruments, local places, and discipline jargon are guessed phonetically. A wrong guess reads like a real word, so skimming does not catch it; checking against the audio does.
Sentence boundaries
Punctuation is inferred, and a moved comma or full stop changes what a participant meant. Fillers and false starts are dropped or kept inconsistently, which matters for clean vs. full verbatim.
Accents, soft speakers, poor audio
Draft quality falls fastest exactly where research interviews are hardest: a quiet participant, a phone-in, an echoing room, a second language.

Measured on one real focus group: Zoom captured 96% of the words, yet 43% of speaking turns held a mistake and 13% changed what the participant said, because the errors landed on the study's own vocabulary. Read the case study, every difference counted.

The corrected file comes back in clean or full verbatim, with speakers labeled and crosstalk marked, in the format you choose. It is the file that goes into the analysis software without anyone on the research team having to listen to the recording again.

What to send

  1. 1

    Send the transcript and the recording

    Zoom saves the transcript as VTT or TXT; a cloud recording usually includes an audio-only M4A, the smallest upload, and a local recording gives you an M4A or MP4. Otter exports DOCX or TXT. Teams saves the transcript as VTT or DOCX and the recording as MP4 in OneDrive or SharePoint. Dragon and other dictation tools give you a DOCX or TXT. Any common audio or video format works: MP3, WAV, M4A, FLAC, MP4.

  2. 2

    Upload both through your Landmark account

    Files go through the HIPAA-compliant platform, never email. Tell us you want the transcript corrected, choose clean or full verbatim, and add speaker identification or timestamps if your analysis needs them. We quote the project before any work starts.

  3. 3

    Or send only the recording

    If you would rather start fresh, send the recording alone and we transcribe it at the standard rate. The AI transcript is not needed for that, and it does not change the price or the turnaround.

Still recording? Headsets, one speaker at a time, and a quiet room do more for the finished transcript than any correction step. Step-by-step upload instructions are in the help center.

When your Zoom transcript is enough

If you only need to find a passage, remember what was discussed, or write up meeting notes, the auto-generated transcript may be all you need — it is searchable and costs you nothing extra. Have it corrected, or the recording transcribed, when the words themselves are the data: coding in NVivo, ATLAS.ti, Dedoose, or MAXQDA, quoting participants in a paper, or documenting procedures for an IRB. Those uses need every word attributed to the right speaker, and that is what the quote pays for.

What you get, and what it costs

Review and correction of an existing transcript is quoted per project from the transcript and the recording. Transcription from the recording is the standard rate: clean-verbatim one-on-one interviews start at $1.99 per audio minute; Plus (focus groups, three or more speakers, full verbatim) starts at $2.69 per audio minute. HIPAA Safe Harbor de-identification and a dedicated project manager are included with transcription. Speaker identification and timestamps are optional add-ons priced with your quote. Standard turnaround is 3–5 business days; rush (24–48 hours) carries an additional charge based on project size and current capacity.

Questions researchers ask

Can you review and correct an AI-generated transcript from Zoom, Otter, Teams, Dragon, or another tool?
Yes. Send the transcript and the recording. A proofreader corrects the transcript word by word against the audio: wrong words, missed speech, speaker labels, sentence boundaries, and the names and terminology the tool guessed. You get back a transcript you can code and quote, delivered in the format you choose.
What does transcript review and correction cost?
It is quoted per project once we have seen the transcript and the recording. The quote depends on the length of the recording, how accurate the draft is, the number of speakers, and the verbatim level you need. Speaker identification and timestamps are optional add-ons at an additional charge, and rush turnaround carries an additional charge based on project size and capacity. You see the rate before any work starts.
Is correcting our transcript cheaper than having you transcribe from the recording?
Not always. Correction is quoted from the draft itself: a transcript that is mostly right takes a proofreader less time than one that has to be retyped in long stretches. If the draft is poor, transcribing from the recording at the standard rate, from $1.99 per audio minute for one-on-one interviews and from $2.69 per audio minute for focus groups and recordings with three or more speakers, can cost the same or less, and we will say so in the quote.
What should I send from Zoom, Otter, or Teams?
Both files: the transcript and the recording. Zoom saves the transcript as a VTT or TXT file and cloud recordings usually include an audio-only M4A, the smallest upload; a local recording gives you an M4A or MP4. Otter exports DOCX or TXT. Teams saves the transcript as VTT or DOCX and the recording as MP4 in OneDrive or SharePoint. Upload both through your Landmark account. If you would rather have us transcribe from scratch, send only the recording.
Does it cost more because the interview was recorded on Zoom or Teams?
No. The recording platform does not change the rate. What sets the rate for transcription is the number of speakers and the verbatim level: one-on-one interviews in clean verbatim are the Standard rate; focus groups, three-plus-speaker recordings, and full verbatim are the Plus rate. Speaker identification and timestamps are optional add-ons priced with your quote, and rush turnaround carries an additional charge based on project size and capacity.
Will the corrected transcript work in NVivo, ATLAS.ti, Dedoose, or MAXQDA?
Yes. Transcripts are delivered in the format you choose and import into the major qualitative analysis packages. If you add timestamps, every speaker change or fixed interval is marked in [hh:mm:ss] so you can jump from a coded passage back to the audio.
Is a Zoom recording of a research interview handled under HIPAA?
On our side, yes: recordings and transcripts are uploaded and stored on Landmark’s HIPAA-compliant platform, handled only by U.S.-based staff, and HIPAA Safe Harbor de-identification is included at no extra charge with transcription. Business Associate Agreements are available for institutions that need one. We cannot speak to how Zoom handles data inside its own service; check your institution’s Zoom agreement and IRB protocol for recording and transcript settings.