SRT to Word Guides

Speakers and scripts

Interview SRT to Word

If the preview never changes speaker, the SRT has no labels. Add [Alex] at the start of the cue, spelled the same way every time, then download.

Do this

All guides
  • Show timestamps On for fact-check, off for the guest
  • Merge same speaker On, after labels exist
  1. Search the SRT for a label at the start of the cue The forms that count are [Alex], Alex:, or a WebVTT <v Alex> tag. A name in the middle of a sentence does not. Zoom files that already have names belong on the Zoom page, not this one.
  2. If there is no label, add one or label later in Word Adding [Host] and [Guest] now is faster when turns are obvious. If they are not, turn merge off, convert, and name people while you listen. Do not guess.
  3. Use one spelling per person Alex and Alexander stay two speakers. The merger compares the label text.
  4. Drop the file with merge on and timestamps on Consecutive cues with the same label, within about 2.5 seconds, become one paragraph under one bold name.
  5. Download interview-timed.docx only after the preview changes voice If the bold name never changes, go back and add labels. Then download the guest copy with timestamps off. Keep the SRT.

Timestamps stay on for the fact-check script. The bold name in the preview must change when the voice changes. Download interview-timed.docx, then a timestamps-off copy for the guest.

A script is not a cue list with the numbers removed

Interview and podcast SRTs are usually captioned for a player: short lines, wraps where the screen ran out, and — if you are lucky — a speaker name. A script reader wants the opposite shape. One person’s turn is a paragraph. The next person’s turn starts on its own line with a name. Cue indexes and arrow lines are gone. Whether the time range stays depends on who will read the file, not on the fact that it began as subtitles.

The converter can build that shape only from labels that are already in the text. It does not listen to the audio, and it does not guess that a lower voice is a second guest. If you need diarization, do it in the tool that made the SRT, or type the names. After the names exist, the Word file is the right place to fix wording.

Labels that become a speaker line

These forms at the start of a cue are read as the speaker. The name is printed in bold. The rest of the cue is the line they said.

  • [Alex] We can cut the cold open.
  • Alex: We can cut the cold open.
  • WebVTT <v Alex>We can cut the cold open.

Brackets are the safest if a guest’s name could be confused with a word. A colon works for ordinary names. The match allows letters, spaces, apostrophes, and hyphens, and it is only attempted once, at the beginning. A sentence like “Note: the bridge is noisy” can be misread as speaker “Note” if it is the first thing in the cue. That is rare in an interview SRT and common in production notes pasted into the same file. Keep notes out of the cue list, or drop them in Word after you see the preview.

A name mentioned mid-sentence stays in the sentence. “Ask Sam about the bridge” does not create a speaker named Sam. A second label on line 2 of the same cue also stays in the text. If your captioner put both people in one cue, split the cue in a subtitle editor or accept a manual break in Word.

Consecutive cues with the same label merge when Merge same speaker is on and the gap between them is about 2.5 seconds or less. Alex’s three caption lines become one paragraph under one bold “Alex”. Sam’s reply starts a new paragraph because the label changed. That is the script. Turn merge off only when you need every caption line preserved as its own block — for example a shot-by-shot paper edit. For a read of the conversation, leave it on.

Unlabeled audio becomes one voice

Auto-captions, a quick Premiere export, and plenty of podcast SRTs have no names at all. Every cue’s speaker is empty. Empty matches empty, so merge treats the whole conversation as one speaker and joins anything within the gap. The preview looks fluent and wrong: a host question and a guest answer in the same paragraph, no cue that the turn changed.

You have two honest options. Add [Host] and [Guest] — or real names — to the start of each cue, then convert with merge on. Or turn merge off, convert, and apply names in Word while you listen. The first option is faster when the turns are already obvious in the caption file. The second is faster when the SRT is a single stream and you will be listening anyway. Guessing names without listening is how the wrong person gets quoted.

Whisper-based interviews often arrive unlabeled and with extra junk lines. Clean the junk before you spend time on names, or you will label a hallucination as a guest. Whisper SRT to Word is that pass. This page assumes the words are mostly real and the missing piece is who said them.

Cross-talk, interruptions, and “both”

Real interviews overlap. SRT can store two cues at the same time. The document lists them in file order, each with its own range if timestamps are on. It does not draw a split screen. If the captioner stuffed both voices into one cue separated by a hyphen, you get one paragraph. Split it if the quote will be attributed.

A label such as [Both] or [Crosstalk] is worth adding when you do not want to pretend one person owns the line. It is still just a speaker prefix. The converter will not detect overlap from the clocks and invent that label for you. Overlapping clocks are a timing topic as much as a script topic; if the ranges themselves are nonsense, fix them before you trust the script as a record. Fix SRT timing is the list of offset and drift cases. Most interview files are not frame-rate problems. They are unlabeled problems.

Timestamps on the record, off the leave-behind

Keep Show timestamps on for the copy you fact-check. A quote that looks too sharp is easier to find at 34:12 than by scrubbing a 90-minute tape with a clean script. Hide timestamps on the copy you send to a guest for approval or paste into show notes, so the page is about the conversation. Download both. Do not delete the SRT after the clean export. The habit is written up in SRT without timestamps.

Zoom recordings are the exception that already has names. The cloud transcript uses WebVTT voice tags or plain “Name / timecode / text” blocks. You usually should not re-label them and you should not re-transcribe the MP4. Start at Zoom transcript to Word for that file shape. Use this page when the interview was captioned as SubRip — a recorder app, an editor export, Descript or another tool that handed you .srt, or a podcast editor’s caption track.

A short pass before you hit download

  1. Search the SRT for your speaker pattern. If you find none, decide whether to add labels now or in Word.
  2. Make the spelling of each name identical. “Alex” and “Alexander” will not merge into one speaker, because the labels differ. Pick one.
  3. Drop the file on the converter and read the preview for one full turn from each person. The bold name should change when the voice changes.
  4. Download the timed .docx for QC. Download the clean one for anyone who should not see cue machinery.

If the preview shows the right names and the wrong music tags or repeated lines, go back to the SRT for one deletion pass. Word is a good place to fix grammar. It is a tedious place to delete the same “[Music]” cue forty times after it has been merged into surrounding speech.

Related workflows

Download the .docx

Timestamps stay on for the fact-check script. The bold name in the preview must change when the voice changes. Download interview-timed.docx, then a timestamps-off copy for the guest.

Speaker questions

How a name gets into the document, and why missing names collapse the interview.

Which speaker formats does the converter recognize?

A label at the start of the cue: [Name], Name:, or a WebVTT <v Name> tag. The name becomes a bold speaker line. A label in the middle of a sentence is ordinary text. Only the start of the cue is read as the speaker.

Why did my two guests become one paragraph?

Unlabeled cues count as the same speaker. With Merge same speaker on, lines within about 2.5 seconds join. If the SRT never marks Alex and Sam, the document cannot tell them apart. Add prefixes, or turn merge off and label the paragraphs in Word.

Is this the same as a Zoom transcript?

Same converter, different source shape. Zoom cloud files are usually WebVTT or speaker-plus-timecode blocks and already carry names. This page is for interview and podcast SRTs where you often have to add the names yourself. Use Zoom transcript to Word when the file came from Zoom.