SRT to Word Guides

Otter → Word

Otter Transcript to Word

Otter’s summary is not the transcript. Convert the SRT, or the speaker-and-time text, and ignore the recap unless you paste it in on purpose.

Do this

All guides
  • Show timestamps On for a record of the conversation
  • Merge same speaker On, so one person’s short lines join
  1. Prefer the SRT if Otter offers one An SRT has --> lines and parses as cues. A conversation export that is already .docx is finished — open it. Do not upload that .docx here; this tool reads caption text, not Word files.
  2. Separate the transcript from the recap Outline, action items, and the AI summary are not timed cues. Pasting them above the transcript makes the preview fail or pulls junk into the first paragraph. Copy the spoken part only.
  3. Match a shape the parser knows Works: an SRT or VTT; a speaker line, then a clock such as 00:12:03, then the sentence; or a line like 00:12:03 Alex: sentence. A free-form paragraph with the time buried in the middle often produces no cues.
  4. Turn merge on and read the bold names Otter splits one answer across several short lines. Merge joins the same speaker when the gap is about 2.5 seconds. If two people share a spelling, they stay one speaker.
  5. Download otter.docx and keep Otter’s export The .docx is for notes and comments. Otter is still the place to correct a misheard name if you will share the conversation from there again.

Drop an Otter .srt, or paste blocks shaped as a speaker line, a timecode line, then the sentence. If the preview says no cues were found, you pasted the summary or a shape this parser does not read. Download otter.docx only after speakers show up.

Otter hands you a conversation and a recap

An Otter conversation often leaves the account in more than one shape: an SRT, a text or Word export of the transcript, and a summary with action items that Otter wrote after the fact. “Otter to Word” searches mix those up. The summary is useful and it is not timed speech. If you paste the whole page, the converter looks for cues and either finds none or treats a heading as a speaker. Delete the recap before you paste, or do not paste it.

If Otter’s own export is already a .docx, that file is the Word document. This site does not open Word files, and running the text of a finished document back through a caption parser can split sentences on clocks that were only there as labels. Use that export as-is. Come here when the file you kept is an .srt or a plain transcript whose speakers and clocks are still in the text.

Shapes that become paragraphs

An SRT or WebVTT with --> lines is the reliable drop. Plain text also works when it matches one of the blocks the parser already uses for meeting transcripts: a speaker line, a clock on the next line (00:12:03 or with milliseconds), then the sentence; or one line that starts with the clock and then Alex: plus the words. Anything looser — a name, a time in parentheses in the middle of a paragraph, a bullet from the summary — is not a cue. The preview will say so. Do not download that empty state and call it a transcript.

Merge same speaker joins Otter’s short lines when the same label repeats within about 2.5 seconds. That is the difference between a caption dump and notes. It will also join two people if Otter spelled them the same, or if the label was missing. Read the bold name in the preview before you download otter.docx.

Where this is not the Zoom page

Zoom’s audio transcript is downloaded from the cloud recording. Otter recorded the meeting because someone invited a bot or uploaded audio. The clocks, the speaker names, and the summary are Otter’s. Use Zoom transcript to Word for a zoom.us VTT, and Teams transcript to Word when the recap is Microsoft’s. Mixing those files in one document is a copy-paste you do in Word after each one has converted, not a mode on this site.

Correct a misheard name in Otter if you will keep sharing the conversation from Otter. Correct the client-facing script in the .docx. Neither edit updates the other.

Related workflows

Download the .docx

Drop an Otter .srt, or paste blocks shaped as a speaker line, a timecode line, then the sentence. If the preview says no cues were found, you pasted the summary or a shape this parser does not read. Download otter.docx only after speakers show up.

Otter export questions

Which part of an Otter export is the transcript, and which part must stay out.

Otter exported a .docx already. Do I convert it?

No. Open Otter’s Word file. This tool does not read .docx uploads. Use it when you have an SRT, a VTT, or plain text with a speaker line and a timecode.

I pasted the whole Otter page and the preview found no cues. Why?

The outline, action items, or AI summary are not timed cues. Delete everything above the spoken transcript, or export SRT and drop that file. A paragraph that mentions a clock in the middle of a sentence is not a cue either.

Will speaker names from Otter stay in Word?

Yes when the text is an SRT with a label, a WebVTT voice tag, or a block this parser knows: speaker, then a clock, then the sentence, or a single line like 00:12:03 Alex: sentence. A name with no clock on its own line may be treated as body text.

Is this the same file as a Zoom transcript?

No. Zoom’s cloud audio transcript is its own download. Otter is a different app even when both can emit WebVTT. Use the Zoom page when the file came from zoom.us.