SubtitleMedic

Convert VTT to TXT

Need the words without the timing? Load a .vtt file, keep the timestamps or merge the lines into paragraphs if you like, and download the dialogue as a .txt transcript.

Processed in your browserNo sign-up.vtt

How to convert VTT to TXT

  1. Open your .vtt file

    Drop one WebVTT file into the box above, or click to browse. It’s read by your browser and never uploaded.

  2. Choose the layout

    Leave both switches off for the words alone, one caption per line. Turn on “Keep timestamps” to start each line with its time, or “Merge into paragraphs” for text that reads like prose.

  3. Download the .txt file

    The preview shows the text exactly as it will be saved. The file keeps the original name with the new extension, so webinar.en.vtt becomes webinar.en.txt.

Where .vtt files come from

WebVTT is the caption format of the web, so most .vtt files are downloads from a platform rather than something a person wrote by hand.

Video platforms. YouTube, Vimeo and most course platforms hand out captions as .vtt. Download tools save YouTube’s automatic captions in the same format.

Meeting recordings. Zoom, Microsoft Teams and Google Meet can save what was said as a .vtt file next to the recording. It is a transcript in disguise: every sentence is there, wrapped in timing lines.

Web players. Any site that shows captions on an HTML5 video loads them from a .vtt file.

In all three cases the words are what you want, for notes, a summary, a search or a translation, and the timing gets in the way.

What is removed, and what stays

The header and its blocks. The WEBVTT line, NOTE comments, and STYLE and REGION blocks describe the file to a player. None of it is dialogue, so none of it is written.

Timing lines and cue settings. Lines such as 00:01:02.345 --> 00:01:04.000 align:start position:0% are left out. With “Keep timestamps” on, the start time comes back in front of the words as [00:01:02].

Tags. Voice tags (<v Maria>), class tags (<c.yellow>), italics and the word-by-word timestamp tags of automatic captions are removed. The words inside them stay.

Character references. &amp; becomes & and &lt; becomes <, so the text reads the way it was spoken.

Repeated lines of rolling captions. Automatic captions repeat each line while the next one is being spoken. Read as plain text, a ten-minute video turns into thirty minutes of the same sentences. The tool keeps each line the first time it appears and drops the repeats; the summary counts the captions that were left out.

What stays. Every word, in order, including sound descriptions such as [applause]. The file is saved as UTF-8.

Plain lines, timestamps or paragraphs

Plain lines suit another program: a translator, a summariser, a word counter, or a search through a whole course.

Timestamps let you find the moment again. They are what you want for quoting a speaker, for study notes, and for writing the chapter list of a long video.

Paragraphs are for people. Captions cut speech into pieces of two or three seconds. Merging joins the pieces that follow each other and starts a new paragraph after a pause of more than two seconds, which is usually where the speaker finishes a point. Automatic captions have no punctuation, so paragraphs are the only structure they get.

Tips for a clean transcript

Automatic captions are not proofread. Names and technical words are often wrong. Fix them in the text file, where a search and replace takes a second.

Keep the .vtt file. A text file can’t become captions again, because the times are gone or reduced to a start time.

Need subtitles for a TV or an editor instead? Text is not what they want: use Convert VTT to SRT. If your file is an .srt, Convert SRT to TXT does the same job as this page.

Frequently asked questions

How do I turn a VTT file into plain text?

Open the .vtt file in the tool and download the .txt file. The WEBVTT header, the timing lines, cue settings and every tag are removed, and each caption becomes one line of text.

Why does every sentence appear two or three times in my YouTube captions?

YouTube's automatic captions roll. Each caption shows the previous line again above the new one, so the raw file repeats everything. The tool recognises these files and writes each line once.

Can I keep the timestamps in the transcript?

Yes. Turn on Keep timestamps and every line starts with the moment it is said, written in hours, minutes and seconds in square brackets.

Are speaker names kept?

Only when they are written in the text itself, such as a name followed by a colon. A name stored in a WebVTT voice tag is part of the markup and is removed with the tag.

What happens to characters like &amp; in the file?

WebVTT writes an ampersand and a less-than sign as character references, because it uses angle brackets for tags. The transcript writes the characters themselves, so the text reads normally.

Are my caption files uploaded anywhere?

No. The file is read and converted by JavaScript in this browser tab. It never leaves your device, and the tool keeps working if you go offline after the page has loaded.