Auto Captions — Word-by-Word, Burned In

Drop a video and get it back with captions burned in, one word lit up at a time. The speech is transcribed and timed on your own device, so nothing is uploaded.

💬 One word at a time🔒 No upload
🎬

Drag & drop your video here, or click to browse

Max file size: ~2 GB (memory permitting)

What the tool looks like

Auto Captions tool with a clip loaded: video preview, the Pop, Punch and Clean caption styles with Pop selected, the bottom, middle and top position choices, and the Add captions button
Dropping a clip in and pressing Add captions is the whole workflow: the audio is pulled out, transcribed with a timestamp on every word, grouped into short phrases and burned into the picture without anything being uploaded.

What it produces

Each caption sits on screen for a few words at a time and the word being spoken is lit up as it is said. Below is the same frame of the same clip through each of the three styles.

A frame of captioned video in the Pop style, with the words "word-by-word captions actually" across the lower third and the current word highlighted
Pop · Yellow active word
A frame of captioned video in the Punch style, with the words "word-by-word captions actually" across the lower third and the current word highlighted
Punch · Green, heavier outline
A frame of captioned video in the Clean style, with the words "word-by-word captions actually" across the lower third and the current word highlighted
Clean · No highlight

How it works

  1. 1
    Drop the video in

    Any common format. The audio is pulled out on your own machine — nothing is uploaded at any point.

  2. 2
    Press Add captions

    The speech is transcribed with a timestamp on every word, the words are grouped into short readable phrases, and the captions are burned into the picture.

  3. 3
    Try another look if you want

    The transcript is kept, so switching style or position only re-renders the video. That takes seconds instead of transcribing again.

What to expect from it

Speech recognition running on your own hardware is a real trade. It costs nothing, uploads nothing and has no per-minute limit, and in exchange it is slower than a server and makes the mistakes any speech model makes on accents, crosstalk and background noise. Read the result before publishing, and use the Subtitle Generator when the exact wording has to be right.

First runAbout 200 MB of model downloads once, then stays in your browser cache.
SpeedRoughly the clip length with WebGPU (Chrome, Edge, Safari 18+); several times that without it.
Best lengthUnder about three minutes. This is built for short-form, and the wait scales with the clip.
Output1080p-safe H.264 MP4 with the original audio copied through untouched.
PrivacyThe audio and the video both stay on the device. There is no server in this workflow at all.

Popular task presets

Best for / not for

Best for

  • Short vertical clips where captions are the difference between someone watching and scrolling past.
  • Talking-head video, podcast cuts, reactions and anything driven by what is being said.
  • Footage that should not be uploaded to a transcription service — client work, internal recordings, anything under NDA.

Not for

  • Precise editorial control over wording. The transcript is used as recognised; correct it in the Subtitle Generator and burn it separately if the exact words matter.
  • Long videos. Transcription runs on your own machine and takes roughly the length of the clip, so an hour of footage is an hour of waiting.
  • Animated or moving caption effects. The words light up in place; they do not bounce, scale or fly in.

When to use word-by-word captions

  • Short-form video where most viewers watch with the sound off — captions are not an accessibility afterthought there, they are the content.
  • Cutting a podcast or interview into clips, where the highlighted word is what keeps a talking head watchable.
  • Anything with a punchline or a number in it, where the timing of the word landing on screen is the point.
  • Recordings you would rather not hand to a cloud transcription service.

Supported formats & limits

Input containersMP4, MOV, WebM, MKV, AVI, FLV, WMV, M4V
Input codecsH.264, H.265, VP8, VP9, MPEG-4, MJPEG. Not AV1 — this WebAssembly build has no software AV1 decoder.
Output containerMP4 (default) — interoperable with iOS, Android, YouTube, Instagram, X
Output codecH.264 video + AAC audio
Max file sizeUp to ~2 GB (limited by browser memory)
Max durationNo hard limit — depends on file size
CostFree for any use. No signup. No watermark.

Auto Captions vs. the usual alternatives

FeatureThis toolVEED (free)Kapwing (free)CapCut Online
Processing modelRuns locally in your browserUpload-based project editorUpload-based project editorUpload-based online editor
File limitsNo upload cap; practical limit is browser memoryPlan-specific upload limitsPlan-specific upload and export limitsFeature- and account-specific limits
Watermark on outputNo watermark addedFree exports include a VEED watermarkFree exports include a Kapwing watermarkStandard edits can be watermark-free; templates/assets may add branding
Signup / accountNo account for toolsWorkspace/account flowWorkspace/account flowCapCut account flow
Works offlineYes after cache, subject to browser supportNoNoNo
Best forPrivate one-step file operationsFull editor, templates, AI toolsCollaboration, templates, AI toolsSocial templates and timeline editing

Vendor plan limits were checked on April 29, 2026 and can change by region, account state, and export option. Verify critical limits on the vendor pricing/help page before relying on them.

Why use this captioner

  • Two actions: drop the file, press the button. The audio extraction, the transcription, the phrase layout and the burn all happen without another click.
  • Real per-word timing, not a phrase split evenly into words — the model is one exported with the cross-attentions that word alignment needs.
  • Changing the style afterwards only re-burns. The transcript is kept, so trying all three looks costs seconds rather than another full pass.
  • Nothing is uploaded, including the audio. The model runs in your browser and is cached after the first use.

Task-focused FAQ

Is my video uploaded anywhere?

No. The speech model is downloaded to your browser and runs there, and the burn-in is FFmpeg compiled to WebAssembly. Neither the audio nor the video leaves the device.

Why is there a download the first time?

The speech model is about 200 MB. Your browser caches it, so the second video you caption starts straight away.

How long does it take?

Roughly the length of the clip on a machine with WebGPU, and noticeably slower without it. A one-minute clip is about a minute; an hour is an hour, which is why this suits short-form.

Can I fix a word it got wrong?

Not in this tool. Use the Auto Subtitle Generator, which gives you an editable transcript and an SRT, then burn that with the Subtitle Burner.

Which languages work?

The model auto-detects and handles the major languages Whisper supports. Accuracy drops on heavy accents and noisy recordings, as it does for every speech model.

Why does the highlighted word sometimes cover two words?

Whisper reports hyphenated compounds as one token, so "word-by-word" lights up together. That is how it was recognised, not a timing error.

Tutorials covering this tool

Frequently Asked Questions

Is my video uploaded anywhere?

No. The speech model is downloaded to your browser and runs there, and the burn-in is FFmpeg compiled to WebAssembly. Neither the audio nor the video leaves the device.

Why is there a download the first time?

The speech model is about 200 MB. Your browser caches it, so the second video you caption starts straight away.

How long does it take?

Roughly the length of the clip on a machine with WebGPU, and noticeably slower without it. A one-minute clip is about a minute; an hour is an hour, which is why this suits short-form.

Can I fix a word it got wrong?

Not in this tool. Use the Auto Subtitle Generator, which gives you an editable transcript and an SRT, then burn that with the Subtitle Burner.

Which languages work?

The model auto-detects and handles the major languages Whisper supports. Accuracy drops on heavy accents and noisy recordings, as it does for every speech model.

Why does the highlighted word sometimes cover two words?

Whisper reports hyphenated compounds as one token, so "word-by-word" lights up together. That is how it was recognised, not a timing error.

Related Tools