A script is not a cue list with the numbers removed
Interview and podcast SRTs are usually captioned for a player: short lines, wraps where the screen ran out,
and — if you are lucky — a speaker name. A script reader wants the opposite shape. One person’s turn is a
paragraph. The next person’s turn starts on its own line with a name. Cue indexes and arrow lines are gone.
Whether the time range stays depends on who will read the file, not on the fact that it began as subtitles.
The converter can build that shape only from labels that are already in the text. It does not listen to the
audio, and it does not guess that a lower voice is a second guest. If you need diarization, do it in the
tool that made the SRT, or type the names. After the names exist, the Word file is the right place to fix
wording.
Labels that become a speaker line
These forms at the start of a cue are read as the speaker. The name is printed in bold. The rest of
the cue is the line they said.
[Alex] We can cut the cold open. Alex: We can cut the cold open. - WebVTT
<v Alex>We can cut the cold open.
Brackets are the safest if a guest’s name could be confused with a word. A colon works for ordinary names.
The match allows letters, spaces, apostrophes, and hyphens, and it is only attempted once, at the beginning.
A sentence like “Note: the bridge is noisy” can be misread as speaker “Note” if it is the first thing in the
cue. That is rare in an interview SRT and common in production notes pasted into the same file. Keep notes
out of the cue list, or drop them in Word after you see the preview.
A name mentioned mid-sentence stays in the sentence. “Ask Sam about the bridge” does not create a speaker
named Sam. A second label on line 2 of the same cue also stays in the text. If your captioner put both
people in one cue, split the cue in a subtitle editor or accept a manual break in Word.
Consecutive cues with the same label merge when Merge same speaker is on and the gap between
them is about 2.5 seconds or less. Alex’s three caption lines become one paragraph under one bold “Alex”.
Sam’s reply starts a new paragraph because the label changed. That is the script. Turn merge off only when
you need every caption line preserved as its own block — for example a shot-by-shot paper edit. For a read
of the conversation, leave it on.
Unlabeled audio becomes one voice
Auto-captions, a quick Premiere export, and plenty of podcast SRTs have no names at all. Every cue’s speaker
is empty. Empty matches empty, so merge treats the whole conversation as one speaker and joins anything
within the gap. The preview looks fluent and wrong: a host question and a guest answer in the same paragraph,
no cue that the turn changed.
You have two honest options. Add [Host] and [Guest] — or real names — to the start
of each cue, then convert with merge on. Or turn merge off, convert, and apply names in Word while you
listen. The first option is faster when the turns are already obvious in the caption file. The second is
faster when the SRT is a single stream and you will be listening anyway. Guessing names without listening
is how the wrong person gets quoted.
Whisper-based interviews often arrive unlabeled and with extra junk lines. Clean the junk before you spend
time on names, or you will label a hallucination as a guest.
Whisper SRT to Word is that pass. This page assumes the words are mostly
real and the missing piece is who said them.
Cross-talk, interruptions, and “both”
Real interviews overlap. SRT can store two cues at the same time. The document lists them in file order,
each with its own range if timestamps are on. It does not draw a split screen. If the captioner stuffed both
voices into one cue separated by a hyphen, you get one paragraph. Split it if the quote will be attributed.
A label such as [Both] or [Crosstalk] is worth adding when you do not want to
pretend one person owns the line. It is still just a speaker prefix. The converter will not detect overlap
from the clocks and invent that label for you. Overlapping clocks are a timing topic as much as a script
topic; if the ranges themselves are nonsense, fix them before you trust the script as a record.
Fix SRT timing is the list of offset and drift cases. Most interview files are
not frame-rate problems. They are unlabeled problems.
Timestamps on the record, off the leave-behind
Keep Show timestamps on for the copy you fact-check. A quote that looks too sharp is easier
to find at 34:12 than by scrubbing a 90-minute tape with a clean script. Hide timestamps on the
copy you send to a guest for approval or paste into show notes, so the page is about the conversation.
Download both. Do not delete the SRT after the clean export. The habit is written up in
SRT without timestamps.
Zoom recordings are the exception that already has names. The cloud transcript uses WebVTT voice tags or
plain “Name / timecode / text” blocks. You usually should not re-label them and you should not re-transcribe
the MP4. Start at Zoom transcript to Word for that file shape. Use
this page when the interview was captioned as SubRip — a recorder app, an editor export, Descript or
another tool that handed you .srt, or a podcast editor’s caption track.
A short pass before you hit download
- Search the SRT for your speaker pattern. If you find none, decide whether to add labels now or in Word.
- Make the spelling of each name identical. “Alex” and “Alexander” will not merge into one speaker, because the labels differ. Pick one.
- Drop the file on the converter and read the preview for one full turn from each person. The bold name should change when the voice changes.
- Download the timed
.docx for QC. Download the clean one for anyone who should not see cue machinery.
If the preview shows the right names and the wrong music tags or repeated lines, go back to the SRT for one
deletion pass. Word is a good place to fix grammar. It is a tedious place to delete the same “[Music]” cue
forty times after it has been merged into surrounding speech.
Related workflows
Download the .docx
Timestamps stay on for the fact-check script. The bold name in the preview must change when the voice changes. Download interview-timed.docx, then a timestamps-off copy for the guest.