In Sync at 0:00, Half a Second Late by 2:00
A growing sync gap is a different fault from a constant offset. Measure at two points, then conform the frame rate, or rebuild a track whose header lied.
The usual sequence goes like this. You notice the lips are off near the end of the video. You drag the audio track back 400 ms, the ending snaps into place, and now the opening is wrong. You drag it back 200 ms and neither end is right. Somebody in a forum thread tells you to install HandBrake.
That advice is not wrong so much as untargeted, because two different faults produce the same complaint. One is a fixed offset: the sound was always late by the same amount, and a timeline nudge cures it. The other is drift, where the gap starts at zero and grows because the video's frames were never written on a regular clock. Nudging cannot cure drift — there is no single number to nudge by.
What follows is the job in order, with a way to check each stage before moving on. Almost everyone who loses an afternoon to this starts at stage five.
Stage 1 — Measure at two points, never at one
Find a moment where sound and picture are provably locked: a hand clap, a hard p on a visible mouth, a key press that makes a character appear on screen. You need one near the start — 0:10, after the player has settled — and one in the last tenth of the runtime.
At each point write down two things: which is ahead, sound or picture, and roughly by how much. Milliseconds are not required. Human tolerance here is lopsided and has been measured for broadcast: sound arriving after the picture tends to go unnoticed up to roughly 125 ms, while sound arriving before it gets caught at around 45 ms. A spoken syllable runs about 0.2 s, so "the mouth finishes half a syllable early" is a perfectly good unit.
If you want something firmer than your eyes, you have to build the measurement out of two tools, because no single one on this site compares the two streams. Frame Extractor is the picture half: it scrubs in 0.01 s steps and prints the time to a tenth of a second, so you can pin the frame where the hands actually meet. It has no transport controls and never plays — it is a seek-only scrubber, and there is no sound in it at all. For the audio half, load the same file into Silence Remover and run only the detect pass. It runs silencedetect and lists every silent span as start–end, also to a tenth of a second; drop the minimum duration to 0.1 s so a short gap registers, and raise the threshold towards −40 dB if a quiet room is not reading as silent. The silence_end immediately before your clap is the moment the clap was heard. Subtract one reading from the other.
Be honest about what that gets you: 0.1 s, which is coarser than the 45 ms you just read about. It is plenty to tell drift from offset and to say which stage you belong in. It is not enough to dial a nudge with, and it only works where the clap sits in silence rather than in the middle of continuous dialogue.
Checkpoint: two readings, each with a direction. One reading is not a diagnosis.
You may want to skip all this and read the file's properties instead. You cannot, not here. Video Metadata reports what a browser hands over — name, size, MIME type, last-modified date, duration and pixel dimensions — and that is the complete list. No frame rate, no codec, no per-frame timestamps, because the <video> element does not expose them and no in-tab tool can invent them. Variable frame timing is precisely what that panel cannot show you, which is why the diagnosis has to be behavioural.
Stage 2 — Read the two numbers before touching the file
| At 0:10 | Near the end | What it is | Fix |
|---|---|---|---|
| ~200 ms, audio late | ~200 ms, audio late | Fixed offset | Shift the track (stage 3) |
| In sync, or too small to see | 0.3–2 s and growing | Variable frame timing | Conform to a constant rate (stages 4–6) |
| Already ~1 s | 20 s or more | Sample rate mislabel | Rebuild the track (stage 7) |
| In sync | Under a second, on a take of an hour | Two devices' clocks diverging | Stretch the external track (stage 7) — not a frame-rate fault |
Row three is separated from row two by sheer size. A sample-rate mislabel — 48 kHz material written into a file that claims 44.1 kHz — plays 48000 ÷ 44100 = 1.088 times too slow. That is about 5.3 seconds of drift per minute, and it drags pitch down roughly a semitone and a half, so voices sound faintly drunk. Variable-frame-rate drift is an order of magnitude smaller: half a second across two minutes is 0.4%, about 0.25 s per minute. If your file is losing whole seconds every minute, stop reading frame-rate advice and go to stage 7.
Row four is the trap, because it is separated from row two by size in the other direction. Two independent recording clocks — a camera and a separate audio recorder — differ by tens to a few hundred parts per million, which is a third of a second to a couple of seconds across a whole hour. It is smaller than the variable-frame-rate case, so it reads like a mild version of row two, and readers conform the frame rate and change nothing, because the picture timing was never at fault. The tell is the shape of the take: dual-system drift needs two devices and a long continuous roll. That whole family — separate recorder, clapperboard, waveform matching, timecode — is covered in Syncing External Audio with Video Footage. This article is about a single file whose own frames were written on an irregular clock.
Checkpoint: you can name the fault and point at the two numbers that justify the name.
Stage 3 — If it is a constant offset, stop here
A fixed gap needs no re-encode, and this site is a poor place to fix one. Nothing here nudges an audio track against picture — no tool passes -itsoffset, -adelay or -apad. That is a timeline operation, and Premiere or Resolve does it in one drag.
The browser-shaped exception is audio running late: the sound lands after the picture. Late audio is cured by making the track start earlier, and you can make a track start earlier by cutting the offset off its head.
- Extract Audio → WAV. The MP3 and AAC options re-encode; WAV is PCM, which keeps the next step exact and costs nothing but disk.
- Audio Cutter, start slider set to your measured offset. Both sliders move in 0.1 s steps and read out in tenths, so round your measurement to the nearest tenth and expect a residual of up to ±50 ms — the same order as the 125 ms you can get away with, which is why this works at all. The cut itself is
-ssbefore-iwith-c copy; on PCM that lands on a sample, so the slider, not the stream copy, is what limits you. (On an MP3 you would also inherit roughly 26 ms of packet granularity. Next to a 100 ms slider step it is noise.) - Audio Replacer, video plus the trimmed track.
Audio Replacer runs -c:v copy -map 0:v -map 1:a -shortest. The picture is copied bit for bit — no second encode, no quality loss, and quick, because nothing is being compressed. Two things to know now rather than four sections from now. First, -shortest truncates the output to whichever stream runs out first, and you just made the audio 200 ms shorter than the video: that 200 ms comes off the end. Trim the video to match, or accept a clipped last moment. Second, no tool on this site passes -c:a copy. Every audio stream that goes through one comes back re-encoded to the container's default codec — AAC for MP4. Budget one audio generation per pass, here and in every stage below.
If the audio runs early — sound before picture, the direction ears catch at roughly a third of the delay they forgive in the other direction — the cure is to push it later by prepending silence, and no single tool here does that. If you can produce a silent WAV of the right length by some other means, Audio Joiner will put it in front: it stages every input to PCM and joins with the concat demuxer, so the lead-in lands exactly where you put it. But nothing here will make that silent file for you — Voice Recorder captures room tone, not digital silence. Early audio belongs in a timeline.
Everything below assumes the gap grew.
Stage 4 — Know what produced the file, then decide how much of it to fix
Accumulating drift almost always means the recorder wrote frames whenever it had one rather than on a fixed grid. Phone cameras do this deliberately: in low light the sensor lengthens its exposure and the rate sags from 30 to 24 or lower, then climbs back outdoors. OBS and game capture do it because the game's frame rate is what it is. Anything recorded through the browser's MediaRecorder does it by design — including the Screen Recorder on this site, which asks the display for an ideal 30 fps and writes VP9 + Opus into WebM. "Ideal" is a request, not a contract. This site produces the problem as well as curing it.
Now the part most walkthroughs skip: decide whether you are fixing an excerpt or the whole recording, because stage 5 is a full re-encode inside a browser tab, and a tab has a ceiling — the uploader refuses anything past 2 GB outright, and a phone gives up long before that. The wasm core here is single-threaded (the site loads ffmpeg-core.js, not the multithreaded build), so budget at least the clip's own running time at 1080p and more on a laptop. A ten-minute file is a ten-minute-plus job you cannot leave the tab for.
Fixing an excerpt first. Use Video Trimmer to take a two-minute span with a sync landmark near each end. This is not the deliverable; it is a cheap test that the conform actually works before you commit an hour to the long job. Video Trimmer runs -ss before -i with -c copy — input seek plus stream copy, so it takes seconds and costs no quality. It also lands on the nearest keyframe, which can add a few tens of milliseconds of its own at the head, so measure the trimmed file fresh rather than comparing its readings to ones you took on the original.
If the whole file has to survive. Split it, conform the pieces, put them back. Video Splitter cuts MP4, WebM and MKV with the segment muxer and -c copy, forcing the cut points onto keyframes, so it is near-instant and loses nothing. Run stage 5 on each chunk, then Video Merger to reassemble — it joins with the concat demuxer and re-encodes to H.264 + AAC, which means the merge is a second generation stacked on stage 5's first. Check one seam before you trust the whole run.
Checkpoint: the piece you are about to conform is a few hundred MB at most on a phone, comfortably under a gigabyte on a laptop, and you have re-measured it. Do not treat the uploader's memory warning as evidence either way. It fires at drop time only, and only on phones and low-RAM machines, where the practical ceiling sits below the hard one; on an ordinary desktop the two are the same number, so the warning never appears at any size the uploader accepts, and its absence promises nothing about whether the encode will finish.
Stage 5 — Conform to a constant frame rate
Video FPS Converter offers 24, 25, 30, 50 and 60, and what it runs is -r <rate> on the output. That flag is the whole cure: FFmpeg rewrites every presentation timestamp onto an even grid, dropping a frame where the source ran fast and duplicating one where it ran slow. On a talking head or a screen recording those handful of edits are invisible. The audio is untouched in time, so it stops sliding against a picture that has stopped wandering.
Pick the rate the file was nominally shot at, not the highest offered. A 30 fps phone clip that sagged to 22 in a dim room goes back to 30; PAL-region footage to 25; game capture to 60 if it was asking for 60. Forcing 60 onto 24 fps material does not make it smoother — it duplicates frames and inflates the file for nothing.
Two things before pressing the button. The tool names its output with the input's extension, so a .webm screen recording comes out .webm, re-encoded with libvpx — considerably slower in WebAssembly than H.264. If you want MP4 at the end anyway, run Video Converter first and conform the MP4 second: the conversion carries the wonky timestamps through for -r to straighten out, and the slow VP9 encode never happens. The corollary matters — Video Converter alone does not fix drift. Its command is -i input output and nothing else; it changes container and codec and preserves the original timing. Only -r conforms.
And unlike stage 3's remarry, this re-encodes the picture as well as the audio. Expect the runtime you budgeted above and one generation of loss on both streams.
Checkpoint: run stage 1 again on the output. Same two points. The gap at the end should now match the gap at the start. Both zero means done. Both 200 ms means you converted a drift problem into an offset problem, which is progress — go fix that in a timeline.
Stage 6 — When the preset you need is not on the list
Five presets do not cover NTSC. 29.97 is 30000/1001 and 23.976 is 24000/1001, and forcing NTSC material to a flat 30 introduces its own slow creep of one frame in a thousand — about 3.6 seconds across an hour, which is small enough to survive a two-minute test and ruin a long one.
FFmpeg Generator is the escape hatch, with one honest caveat: none of its nine operations is a frame-rate change — they run from trim and convert through compress, merge and remove audio. Choose Convert Format, copy the command it builds, and insert the rate yourself:
ffmpeg -i input.mp4 -r 30000/1001 input_output.mp4
For film telecined into a 29.97 stream, -vf pullup -r 24000/1001 reverses the pulldown instead of papering over it. Both run in desktop FFmpeg, not in the tab — the generator writes commands, it does not execute them. If you are opening a terminal anyway, do the conversion and the conforming in that one command: one encode beats two, and it beats split-conform-merge outright.
Stage 7 — The drift a constant frame rate will not touch
If stage 2 pointed at row three or row four, -r is irrelevant. The picture timing is fine. Rewriting frame timestamps changes nothing about an audio stream that is being consumed at the wrong speed.
Row three: the header lies about the sample rate. The browser toolkit is not merely imprecise here, it is the wrong shape. Audio Speed & Pitch has a speed slider that steps in increments of 0.05, so the 1.088 you need is unreachable — 1.10 overshoots by about 1%, still six seconds of drift across ten minutes. Worse, that slider is built on an atempo chain, which time-stretches while preserving pitch; the tool's own note says so. A sample-rate mislabel makes the audio both slow and flat by about 1.47 semitones, so even an exact 1.088 would fix the duration and leave the voice a semitone and a half under. The separate pitch control steps in whole semitones from −12 to +12, so 1.47 is out of reach there too. What this fault needs is speed and pitch moving together, which is precisely what that tool is built not to do.
One line in desktop FFmpeg does it, because asetrate re-declares the rate rather than resampling:
ffmpeg -i input.mp4 -vn -af "asetrate=48000,aresample=48000" fixed.wav
Use the rate the material was actually recorded at, not the one the header claims. Then bring fixed.wav back to Audio Replacer, exactly as in stage 3 — same -shortest warning, same container-default re-encode of the audio, picture still copied untouched.
Row four: two clocks. A camera and a separate recorder running off their own crystals diverge by tens to a few hundred ppm; 100 ppm is 0.36 s per hour. There is no flag for this, and -r makes it worse by encouraging you to believe the picture was wrong. The cure is a ratio you derive from a slate at each end of the roll: divide the camera's clap-to-clap interval by the recorder's, and stretch the external track by that. Four decimal places, in a timeline's speed or conform control, or with atempo in desktop FFmpeg. Here — unlike row three — a pitch-preserving stretch is the right tool: even a few hundred ppm is well under a hundredth of a semitone, inaudible either way.
Checkpoint: stage 1 one final time, and confirm the last second of the file still exists.
What each skipped stage costs
Skip stage 1 and you spend an hour on the wrong fault: the nudge that ruins the opening, or the conversion that does nothing because the problem was the sample rate. Skip the magnitude check in stage 2 and you either conform a file that was never variable, paying a full re-encode to change nothing, or you read a dual-recorder take as a frame-rate fault and conform the one stream that was already correct.
Skip the excerpt in stage 4 and you find out whether -r was the right answer after a forty-minute encode instead of a four-minute one. Skip the second measurement in stage 5 and you publish something that is merely differently wrong.
Getting the order wrong — converting after conforming rather than before — breaks nothing. It only makes you pay for the slow encode twice.
Try it yourself — free in your browser
No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.