Transcript Tool

Pull the transcript out of any video

Paste a link or drop in a file and we listen to it, then write down every word that is said — as clean, copyable lines with timestamps. Nothing to install.

Free to use No sign-up Nothing stored

How to get a transcript from any video

A video transcript is the spoken content of a video written out as text. Paste a link or upload a file, and this tool listens to the audio and writes down what is said, line by line, with the moment each line was spoken. You get text you can read, search, quote, translate, or turn into subtitles — without watching the video again.

Two ways to start

  • Paste a video URL. Works with YouTube, TikTok, Facebook, Vimeo, Dailymotion, X and most public video pages. Nothing is installed and you do not need an account.
  • Upload a video file. MP4, MOV, AVI, MKV or WebM, straight from your device — useful for recordings that were never published anywhere: interviews, lectures, meetings, voice notes, footage you shot yourself.

Why this is not just "caption scraping"

Most free transcript tools only fetch subtitles a platform has already published. If a video has no captions, those tools return nothing. This one runs speech recognition on the actual audio, so it works on videos that have never been captioned — including your own uploads.

There is one deliberate shortcut. When a person has written captions for a video, those are used instead, because a human transcript is more accurate than any model. The result tells you which one you got, so you always know whether you are reading a person's work or a machine's.

What you can do with the transcript

  • Jump back into the video. Every timestamp is a link that opens the source video at that exact second — the fastest way to find the moment something was said.
  • Search inside the words. Type into the search box to highlight every occurrence and step through the matches.
  • Read it as prose. Switch from timestamped lines to paragraphs, grouped where the speaker actually paused.
  • Quote it properly. Copy any line as a citation — the quote, the video title, the timestamp and a link straight to that moment.
  • Export it. Plain text, Markdown, CSV, JSON, or subtitle files in .srt and .vtt.

Turn on timestamps before you extract

The Include timestamps switch decides how much you get back. With it off you receive clean running text. With it on, every line carries the time it was spoken, which is what unlocks the jump links, the citation copy, and the subtitle exports. If you plan to caption a video or cite it, switch it on first.

Who uses a video transcript, and why

  • Journalists and researchers quote accurately and link to the exact second a statement was made, instead of paraphrasing from memory.
  • Creators turn a video into a blog post, newsletter, or set of captions without retyping it.
  • Students and note-takers skim an hour-long lecture in a couple of minutes and search it for the part that mattered.
  • Accessibility teams produce .srt or .vtt subtitle files so a video can be watched without sound.
  • SEO and marketing teams publish the transcript alongside a video, so search engines can read content that would otherwise be locked inside a media file.

Getting the best results

Speech recognition is very good, but it is not magic. A few things make a noticeable difference:

  • Clear audio beats loud audio. One person speaking close to a microphone transcribes far better than several people across a room.
  • Heavy background music is the most common cause of missing or garbled lines.
  • Names, brands and technical terms are where machine transcription slips most often. Skim those before publishing.
  • Long videos take longer. The audio is processed at a few times real time, so a short clip returns in seconds and a long one takes a few minutes.

Accuracy, honestly

When the words come from human-written captions, they are as accurate as the person who typed them. When they come from speech recognition, expect the occasional slip — especially with strong accents, crosstalk, or specialist vocabulary. The tool labels which source it used and never invents words: if there is no speech in a video, it says so rather than filling the silence.

Privacy

Your video is read, transcribed, and deleted — nothing about it is kept, and no transcript is stored on our servers. Transcription runs on our own machines rather than being handed to a third-party service, and the text is returned to your browser, where the copying and exporting all happen locally.

Frequently confused: transcript, subtitles, and captions

A transcript is the full text of what was said, meant to be read on its own. Subtitles are that text split into short timed chunks and displayed over the video — that is what an .srt or .vtt file contains. Captions usually mean subtitles that also describe meaningful sound, such as [door slams], for viewers who cannot hear the audio. This tool gives you the transcript, and exports subtitle files from it when timestamps are switched on.

FAQ

Questions, answered

YouTube, TikTok, Facebook, Vimeo, Dailymotion, X and most public video pages — plus any video file you upload from your own device. Because we transcribe the audio rather than look for existing captions, a video does not need subtitles for this to work.
Yes. Speech recognition runs over the actual audio and writes down what is said, so it works on videos that have no captions at all. The one shortcut: when a person has already written captions for a video, we use those, because a human transcript beats any model.
Yes — switch on Include timestamps before extracting. You then get a time against every line, and can export the result as an .srt subtitle file as well as plain text.
No. The audio is read, transcribed, and deleted as soon as the words come back — nothing about your video is kept, and transcription runs on our own machines rather than being sent to a third party. See our Privacy Policy.