EditingIntermediate

Joining AI Clips Into One Video Without a Broken Seam

Merged clips from different AI generators freeze, go silent or judder after clip one. A symptom table, the fix per cause, and the one-pass desktop command.

Applicable Software:Premiere ProDaVinci Resolve

The command every answer hands you is this one:

ffmpeg -f concat -safe 0 -i list.txt -c copy output.mp4

It is the right command for eleven clips off one camera, shot in one session, one codec, one size. It is the wrong command for eleven clips out of two or three different generators, and -c copy is the reason. Video Merger refuses to run it — the comment in its source says why: the concat demuxer with -c copy produces broken output when the inputs have different codecs, profiles, timebases or sample rates, which is the common case for clips grabbed from different sources. So it re-encodes instead, with libx264 -preset veryfast -crf 23 -c:a aac -b:a 192k.

Here is the part that catches people who already knew that. Re-encoding fixes the output. It does not fix the input. The concat demuxer still stitches the files into one virtual stream before the encoder ever sees a frame, and it takes its stream parameters from the first file in the list. Every failure below happens on the read side, upstream of the re-encode. So the order that works is: trim each clip, conform its shape, conform its rate, then merge once.

Find your symptom

What you see What actually happened The one thing to do
Clip 1 plays, then a frozen frame, smeared blocks, or the run just dies The clips are different pixel sizes, and the geometry was fixed by clip 1 Send every clip through Video Resizer at one preset
Everything after clip 1 is garbage, no error Mixed codecs — a VP9 WebM and an H.264 MP4 in the same list Same fix: the resizer re-encodes to H.264 MP4 whatever you feed it
Audio on clip 1, silence for the rest At least one clip has no audio stream, and the layout was built from clip 1 Mute the whole set, add one music bed after the merge
The back half stutters or judders; picture drifts against sound Mixed frame rates Pick one rate and convert every clip to it
A soft or warped frame at each join The generator's first and last frames Trim before merging, never after
Tab reloads, or you get a 0 KB merged.mp4 Memory — every input sits in the WASM filesystem at once Merge in batches, knowing what each batch level costs
Fewer clips in the list than you selected The merger only accepts files whose MIME type starts with video/, and drops the rest silently Check the count in the list header before merging
Clips joined in the wrong order The list is in file-picker order, not filename order Renumber 01_, 02_, or use the ▲▼ buttons

Different sizes, and therefore a dead clip 2

The most common failure, and it survives the re-encode because it happens before the re-encode. The demuxer presents N files to FFmpeg as one continuous stream, and the decode and encode are configured once, from the head of it. When a 1280×720 clip arrives behind a 1920×1080 one, nothing downstream can renegotiate a frame size that is already open: depending on the build you get a freeze, smeared blocks, or a run that dies and leaves nothing. What you never get is an error naming the clip.

One preset, applied to all eleven, whichever preset you pick — that is the whole rule, and Video Resizer is where you apply it. Which target to choose, and whether to pad or crop into it, is a real question with a real answer, but it is a different article's question; here any one target beats no target.

The part worth knowing is what else the resizer does on the way past. It always writes H.264 in MP4 at CRF 23, preset veryfast, AAC 192k, with setsar=1 on the end of the filter chain, regardless of what went in. One pass conforms geometry, codec, container and pixel aspect ratio together. That is why it is worth sending a clip through it even when it is already the right size — it is the only tool here that normalises all four at once. Video Converter fixes container and codec and leaves the geometry alone, so it is not a substitute.

One frame rate for the whole set

The merger passes no -r, so the output is conformed to a single rate taken from the stream the demuxer set up first, and every clip that disagrees has frames duplicated or dropped to fit it. Timestamps are carried through with an offset per segment, so the total duration comes back correct — which is exactly why this one gets missed. The symptom is stutter and judder in the later clips, plus drift between picture and any sound you have, not a file that is the wrong length. Conform everything with Video FPS Converter before the merge, not after.

Do not reach for frame interpolation to smooth the result. minterpolate and optical-flow smoothing synthesise in-between frames by warping what is already on screen, and what is already on screen in a generated clip is precisely what the model renders least stably: hands, faces, lettering. A duplicated or dropped whole frame reads as ordinary footage. A warped finger does not.

Audio on the first clip, then silence

Plenty of generations still come back with no audio stream, and some now come back with one. A list that mixes the two is the classic way to lose sound after the first join: the stream layout is built from the first file, and files that do not match it contribute nothing.

You cannot diagnose this here, and it is worth knowing that before you go looking. Video Metadata reads an HTML <video> element's loadedmetadata event, so it reports file name, size, MIME type, duration and resolution — no tool on this site reports whether an audio stream exists. The workable check is manual: open each clip in a browser tab and watch for the tab's audio indicator. No indicator, no stream.

The cheaper move is to skip diagnosis and normalise downward. Run Video Mute over the whole set so it is uniformly silent — it is -an -c:v copy, instant and costing nothing in quality — merge, then put sound back with Add Music to Video, which maps the track on with -c:v copy so the picture is never encoded a second time.

If some clips carry dialogue you actually need, that does not work, and nothing here will manufacture a matching silent track for the clips that lack one. On a desktop it is one command per silent clip:

ffmpeg -i silent.mp4 -f lavfi -i anullsrc=r=48000:cl=stereo \
  -map 0:v -map 1:a -c:v copy -c:a aac -shortest fixed.mp4

Set r= to the sample rate of the clips that do have sound, and now every file in the list has the same layout.

The soft frame at every join

Generations open and close weakly — a fraction of a second of resolve at the head, a drift or a morph at the tail. Cut both off with Video Trimmer before anything else. It runs -ss before -i with -c copy, so it is instant and lossless, but the stream copy means the start can only land on a keyframe at or before where you left the slider. The end is not constrained the same way: -t stops where you asked.

The failure mode nobody warns you about is what that means on short clips. A five-to-ten-second generation often has exactly one keyframe — its first frame. Ask for a 0.4 s head trim and the cut snaps back to zero, the soft frame survives untouched, and the trimmer looks broken. It is not; there was no keyframe to snap to. When that happens, go to Video Splitter, tick "Precise cuts (re-encode for frame-exact splits)", split at 0.4 s and keep part two. It costs one re-encode, which is the going rate for a frame-accurate cut.

Trim first, not last. After the merge those boundaries are interior to one file, and removing them means a split plus a second merge — a full CRF 23 pass over the whole timeline instead of a lossless snip of five seconds.

The tab dies, or the download is 0 KB

Everything runs in the page, so the ceiling is how much memory the browser gives one tab. The trimmer, resizer and FPS converter go through the shared uploader, which hard-caps at 2048 MB and warns above a device-derived figure below that. The merger does not use that uploader at all: no size check, no warning, and it writes every input into the WASM filesystem before the concat starts, so peak memory is roughly the sum of all inputs plus the output plus decode buffers.

A dozen 1080p clips of a few seconds each is a couple of hundred megabytes and fits comfortably. The same dozen at 4K generally does not. You can merge three or four at a time and then merge the results — but price it first: each extra merge level is another CRF 23 pass over everything inside it, so two levels puts your footage through four generations of compression instead of three. That is the most expensive advice on this page. Read it as the signal to move, not as a free workaround.

Worth knowing: the tools that run through the shared exec helper check for an empty result and raise "The video engine produced no output… this usually means the file uses a codec the in-browser engine cannot read." The merger reads its output directly and skips that check, so a failed concat can hand you a zero-byte merged.mp4 and still say Done. A 0 KB download is a failed merge, not a failed download.

If one specific clip kills every run, suspect AV1. The build has no software AV1 decoder — AV1 input fails outright, and nothing here can convert it. Re-encode that clip to H.264 outside the browser first.

If none of these matched

Do not walk the list from the top. Pair against a known-good instead: find two clips that merge cleanly, then swap one of them for a suspect. A clip that fails against every partner, and also fails alone through Video Converter as a solo re-encode, is bad in itself — a truncated generation, a codec the build cannot read, a broken stream. A clip that merges with some partners and not others is a mismatch, and the table above names which kind.

Video Metadata helps with exactly two of the columns you need — resolution and duration — and not with codec, frame rate or audio, because it reads the browser's own decoder rather than the file header. Resolution alone still catches the most common offender in the list.

The arithmetic that decides it

Two numbers, and they are the honest reason people abandon this halfway.

Passes. Resize is one re-encode, the FPS conversion is a second, the merge is a third. Three generations of CRF 23 stacked on whatever the generator already compressed. It holds up on motion and skin. It stops holding up on fine text, hair and logos.

Hands on keyboard. Three per-clip tools across N clips is 3N upload-encode-download cycles, one file at a time, because there is no batch mode anywhere on this site. Eleven clips is 33 of them, plus renaming to keep the order straight, plus a 32 MB engine that each tool page loads before its first run. The encodes themselves are quick on a few seconds of 1080p; the file shuffling is the afternoon.

The one-pass alternative is the concat filter, not the demuxer — and be exact about what it forgives. Different codecs and containers, yes. Different geometry, no: every segment must share width, height, SAR and pixel format, and a 1280×720 input beside a 1920×1080 one stops with Input link in1:v0 parameters (size 1280x720, SAR 1:1) do not match the corresponding output link parameters. The conform work does not vanish. It moves into the filter graph, per input, where it costs no extra encode:

ffmpeg -i a.mp4 -i b.mp4 -i c.mp4 -filter_complex \
"[0:v]scale=1080:1920:force_original_aspect_ratio=decrease,\
pad=1080:1920:(ow-iw)/2:(oh-ih)/2:black,setsar=1,fps=30[v0];\
[1:v]scale=1080:1920:force_original_aspect_ratio=decrease,\
pad=1080:1920:(ow-iw)/2:(oh-ih)/2:black,setsar=1,fps=30[v1];\
[2:v]scale=1080:1920:force_original_aspect_ratio=decrease,\
pad=1080:1920:(ow-iw)/2:(oh-ih)/2:black,setsar=1,fps=30[v2];\
[v0][v1][v2]concat=n=3:v=1:a=0[out]" \
-map "[out]" -c:v libx264 -preset veryfast -crf 20 out.mp4

Three things scale with the clip count: one -i per clip, one scale/pad/setsar/fps leg per clip, and n= must equal the number of legs you wrote. a=0 is the right choice when the generations are silent — the [0:a][1:a]…a=1 form that gets quoted everywhere fails with Stream specifier [1:a] matches no streams the moment one input has no audio track. Give the silent ones an anullsrc track first if you genuinely need a=1.

Eleven inputs, without typing eleven legs:

clips=(clip*.mp4)          # name them clip01 … clip11 so the glob sorts right
n=${#clips[@]}
leg='scale=1080:1920:force_original_aspect_ratio=decrease,pad=1080:1920:(ow-iw)/2:(oh-ih)/2:black,setsar=1,fps=30'
fc=''; outs=''; ins=()
for i in "${!clips[@]}"; do
  fc+="[$i:v]$leg[v$i];"; outs+="[v$i]"; ins+=(-i "${clips[$i]}")
done
ffmpeg "${ins[@]}" -filter_complex "$fc${outs}concat=n=$n:v=1:a=0[out]" \
  -map "[out]" -c:v libx264 -preset veryfast -crf 20 out.mp4

One warning about where not to get that command. FFmpeg Generator has a Merge task, and it does not write this. It emits the -c copy concat-demuxer version — the exact command this article opened by warning you off — so copy the filter form from here instead. Premiere and DaVinci Resolve arrive at the same place from the other direction: sequence settings conform every import, and you render once.

Try it yourself — free in your browser

No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.

Tags:ai videomergingffmpegconcattroubleshootingworkflow