EditingBeginner

Captioning a One-Hour Video After the Free Minutes Run Out

Browser Whisper has no free-minute cap, but it bills in model downloads, RAM and a Latin-only burn font. Which of the two tools to use, and when to split.

Applicable Software:Premiere ProDaVinci Resolve

Metered captioning dies at the same file every time. Thirty free minutes, sixty, three hundred a month — one lecture, one sermon, one council meeting swallows the whole allowance, and the recording after it costs money. Whisper running inside your own browser has no allowance. It does have a bill, and none of it appears on a pricing page: a one-time model download, wall-clock time that scales with the length of the recording, a few hundred megabytes of RAM held for the duration, and a bundled font that can only draw about half the languages the transcriber can hear.

Two tools here spend that bill, and they spend it very differently. Their names suggest the same thing at two sizes. They are not that.

Two tools, one engine

Both pull transformers.js from jsDelivr the first time you run them, and both fetch ONNX weights from the Hugging Face CDN. After that they diverge on every axis that matters for a long file.

Auto Subtitle Generator Auto Captions
What you get SRT, VTT and TXT, editable line by line before download A burned-in MP4 and nothing else — no transcript, no text export
Model Your pick: Fast ≈ 114 MB, Balanced ≈ 196 MB, Accurate ≈ 558 MB One fixed word-timestamped checkpoint, ≈ 197 MB
Language Auto-detect or 20 named languages, plus a transcribe/translate choice Auto-detect, transcribe. No language or task control exists
Awkward containers Falls back to FFmpeg when the browser refuses the file, so MKV and AVI work Reads frame dimensions from a <video> element first; if your browser cannot play it, the button stays grey with no explanation
Size guard None. It takes whatever you drop on it Warns above your device's practical limit, refuses anything over 2 GB

Translate only ever goes one direction, into English. And the twenty-language dropdown is a convenience, not a ceiling — auto-detect hands Whisper every language it knows, so Greek or Hebrew transcribe fine despite not being listed. Hold on to translate: it becomes the only escape hatch from the font problem below, and it lives on exactly one of these two routes.

Auto Captions loads a fourth model instead of reusing Balanced because word-level alignment is recovered from the model's cross-attentions, and the ordinary Whisper exports do not emit them — ask them for word timestamps and the run dies with Model outputs must contain cross attentions. Both tools run an fp32 encoder with a q4-quantised decoder, because the q8 exports of every Whisper size refuse to build an inference session at all, which would take out the CPU path entirely. That is why the two "base" models land within a megabyte of each other and still cannot share a download.

Each model is cached by the browser after its first fetch, so your second file costs nothing but time. The caches are separate, though: moving from Balanced to Accurate is a fresh 558 MB, and Auto Captions pulls its 197 MB even if base is already sitting there.

Which leads to the fact that settles this for most long recordings. Auto Captions cannot be given an SRT. It accepts a video and always transcribes from scratch. Run Auto Subtitle Generator for the transcript and then Auto Captions for the burned-in copy, and you have paid for the transcription twice, on two different checkpoints, with an extra 197 MB in between. If you want both the text and the pixels, the route is Auto Subtitle Generator → Subtitle Burner: one download, one transcription, and you keep the SRT.

Which model

Fast (tiny). One speaker, a headset or lavalier mic, a language the model has heard a great deal of, and you intend to read every line anyway. It is also the only sensible choice if your browser has no WebGPU and you want the transcript today rather than tonight. Expect it to mangle proper nouns and any term narrower than general vocabulary.

Balanced (base). The default, and correctly so for a recorded talk in a decent room. It is the point where you stop fixing a word every other line and start fixing one every other minute.

Accurate (small). Accents the tiny model stumbles on, crosstalk, air conditioning, or a subject with real vocabulary — medical, legal, a codebase, a liturgy. The 558 MB is a one-time cost; the slower inference is charged on every run. Pick it when the alternative is an hour of correcting by hand.

Auto Captions. Short-form video, and essentially nothing else. It groups words into lines of at most four words or thirty characters, breaking early on a full stop or a pause longer than 0.6 seconds, and lights one word at a time in two of its three presets — the third, Clean, drops the highlight and leaves a plain white caption. For a sixty-minute lecture that is the wrong shape of caption and an enormous amount of machinery to produce it: an hour of speech at a normal 140 words a minute becomes roughly 8,400 subtitle events, because the highlight is built as one dialogue line per word rather than with karaoke tags.

What the hour actually costs

Transcription runs in fixed 30-second windows with a 5-second stride on each side, one window after another. Nothing about a sixty-minute file differs in kind from a sixty-second one — there is no ceiling in the code and no quota to hit, and minute 55 costs what minute 5 did. The queue just gets longer.

Which is why the honest answer to "how long" is a method rather than a number. Cut five minutes off the front of your recording, run it on the model and the machine you actually intend to use, and multiply by twelve. That figure beats any benchmark quoted at you, because the spread between a laptop with WebGPU and the same laptop without it is wider than the spread between models. The tool's own estimate — roughly the length of the clip on Fast without WebGPU, longer on the others — appears in an amber box in Auto Subtitle Generator and nowhere in Auto Captions.

The second lever is yours: Fast is 39 million parameters, Balanced 74 million, Accurate 244 million. Inference time does not track parameter count exactly, but closely enough that Fast-versus-Accurate is the largest decision on the page — larger than the container, the language, or anything in the style controls.

Memory is arithmetic rather than opinion. An hour of audio at 16 kHz mono in 32-bit floats is 3,600 × 16,000 × 4 bytes — about 230 MB, one contiguous array, held for the whole run alongside the model. Before that array exists the browser decodes the file with decodeAudioData; the tool asks for a 16 kHz context, but the code carries an explicit fallback for browsers that ignore the request and decode at the file's native rate, and an hour of 48 kHz stereo is roughly 1.4 GB before it is downsampled. That spike lands before a byte of model has loaded. The file itself is read whole via arrayBuffer() first, so its size sits on the heap as well.

Nothing warns you about any of this on the recommended route. Auto Subtitle Generator has its own dropzone with no size check at all — no device lookup, no amber box, nothing. The 2 GB hard refusal and the device-limit warning live in the shared uploader, which Auto Captions, Subtitle Burner and Extract Audio use and the generator does not. On a phone that means the 230 MB buffer plus a model plus the decode spike is simply a reclaimed tab, with no warning first. It also means that above 2 GB nothing on this site will accept the file — Extract Audio refuses it before it starts — so a recording that large has to be cut down somewhere else first.

Cut the hour up before you spend it

An hour in one run is an hour you can lose in one reload. Nothing touches disk until you download, the tab has to stay open, and a reclaimed tab costs the whole afternoon. Splitting first turns that into a fifteen-minute loss.

  1. Split. Video Splitter in "every N seconds" mode, with 900 for fifteen-minute parts. It stream-copies by default, so it is near-instant and lossless; the trade is that copy-mode cuts land on the nearest keyframe, so your parts will not be exactly 15:00 each. If you want exact boundaries, run the recording through Extract Audio to MP3 first — about 85 MB an hour at its VBR setting — and cut that with Audio Cutter, which also copies rather than re-encodes.
  2. Transcribe each part. The model downloads once and is reused for all four.
  3. Shift each SRT. Part two begins at 15:00 in the finished video, so its subtitles need +900000 ms in SRT Offset. Type that into the number box rather than dragging the slider — the slider stops at ±5 seconds, the box takes any integer. If you split with stream copy, use the cumulative real durations of the earlier parts instead of 15:00 × n.
  4. Paste them together. FFmpeg's SRT reader keys off the timestamps, so repeated index numbers in a stitched file are harmless; renumber only if the player you are targeting is fussy about them.

The cost is real and worth knowing in advance: Whisper sees no context across a cut, so a sentence spanning a boundary can break in two and the punctuation at the seam gets strange. On a lecture that is one awkward line every fifteen minutes. The alternative, on a reload, is all of it.

The burn is a separate, larger job

Getting an SRT is the hard part intellectually. Getting it onto the pixels is the part that fails silently, and there is one reason for that: this FFmpeg build has no fontconfig. Ask libass to draw with a font name and it reports fontconfig is not available, ignoring font option, finds nothing, and draws absolutely nothing — while still re-encoding the video and handing you a file that looks like success. That is why Subtitle Burner writes a Noto Sans TTF into the virtual filesystem and passes subtitles=<file>:fontsdir=/fonts explicitly.

The font it hands over is a Latin subset: 226 codepoints, ASCII plus Latin-1 plus a handful of punctuation. Outside Latin-1 it carries exactly three letters — ı, Œ and œ — which tells you how thin it is, since that means French is safe and Turkish is not, despite Turkish being the reason the dotless ı is in there at all. The subset covers English, Spanish, French, German, Italian, Portuguese, Dutch, Swedish and Indonesian: nine of the twenty languages in the dropdown. Polish ł and ą, Turkish ğ and ş, every Cyrillic letter, Vietnamese tone marks, Arabic, Hindi, Thai, Japanese, Korean and Chinese are all absent.

What you get in that case is not an invisible caption. libass reports the glyph as not found, falls back to the font's own .notdef, and draws a row of empty rectangles — tofu — where the non-Latin characters were, with any Latin words on the same line rendering perfectly. Visibly broken rather than invisibly absent; the invisible failure is the other one, the fontconfig one. The escape hatch is narrow and specific: set the task in Auto Subtitle Generator to translate, which gives you English, and burn that SRT in Subtitle Burner. Auto Captions uses the same Latin-only font — the same two Noto files, with Fontname: Noto Sans hardcoded into the ASS it generates — and, having no language or task control, offers no way around it. For a non-Latin language, Auto Captions is out entirely.

Three more things before you commit an hour of encoding:

  • The burn is a full re-encode, every time. There is no stream-copy path for video; even a burn that draws nothing re-encodes the lot. Subtitle Burner passes no -c:v, -preset or -crf at all, so it takes libx264's defaults — preset medium, CRF 23 — on a single-threaded WebAssembly core. Auto Captions burns at -preset veryfast -crf 20. On an hour of 1080p the burn can comfortably outlast the transcription, and the burner is the slower of the two. Audio is stream-copied, which is the cheap half and survives fine, Opus into MP4 included.
  • The video has to be decodable. This is where AV1 bites, and only here: there is no software AV1 decoder in the build. Transcription is unaffected, because the audio path never opens a video decoder — the browser decodes the audio track directly, and the FFmpeg fallback runs with -vn. So an AV1 recording transcribes perfectly and then refuses to burn. The same pattern applies to an MKV in Auto Captions, which needs the browser to play the file before it will start. Video Converter solves both, at the price of one more full pass over the hour.
  • An ASS file's own styling wins. Load one and the size and position dropdowns grey out by design; only the font family is forced. If you ran your SRT through Subtitle Converter to make an ASS "because it's more advanced", you have locked yourself into that converter's hardcoded 60 px bottom-centre style. Burn the SRT directly and you keep the 18/22/28/36 px choice and top/middle/bottom placement.

Re-styling, at least, is cheaper than re-transcribing. Auto Captions keeps the word list in memory after a successful run, so pressing the button again with a different style skips transcription entirely. On a thirty-second clip that genuinely is seconds. On an hour it is still the full single-threaded H.264 encode described above — cheap only relative to transcribing again. And the word list is memory only: reload, reset, or swap the file and the expensive part restarts from zero.

When the answer is none of these

  • Your language is not Latin-1, the captions must be burned in, and English will not do. Take the SRT and burn it somewhere with real font handling. Premiere Pro and DaVinci Resolve both draw from the system font stack, so they will render whatever your machine can render.
  • You have a stack of hour-long files. Whisper running natively is several times faster than WebAssembly, handles every container, and does not lose its work when a tab is reclaimed. A whisper.cpp binary, or one of the desktop GUI wrappers built on it, is the low-friction option; faster-whisper if you are willing to touch Python. Each costs you a real install and a model download, and each buys you batch processing and no tab to babysit. One file is a good trade; twenty is not.
  • You need a verbatim, speaker-labelled transcript. Whisper does no diarisation here, and no in-browser tool in this set adds it.

And one case that is not "none of these" at all. If the captions only need to appear during playback, do not burn anything. Now that the burn is a known quantity — a full single-threaded encode of the entire hour, plausibly longer than the transcription that produced the text — the argument for keeping subtitles soft is mostly arithmetic. Keep the SRT, check it against the video in Subtitle Preview, and hand the file to YouTube, Vimeo or VLC alongside the video. You skip the encode completely, the text stays searchable and correctable, and the font problem vanishes: the player brings its own.

Try it yourself — free in your browser

No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.

Tags:subtitleswhisperauto captionstranscriptionsubtitle burningwebassembly