Duck the Music Under the Voice

Add a music bed that drops automatically whenever someone speaks and comes back up when they stop. Sidechain compression, in the browser.

🔒 100% Private — no upload🎬 MP4, WebM, MOV, AVI⚙️ No install needed

Ducking is the music getting quieter by itself whenever someone speaks, and coming back up in the gaps. Only the music is compressed: the voice is the trigger and is mixed back at its own level, and the picture is copied through without re-encoding.

🎬

Drag & drop your video here, or click to browse

A voice-over on its own works too. Max file size: ~2 GB (memory permitting)

🎵

Choose an MP3, WAV, M4A or OGG

How to Use — Duck the Music Under the Voice

1

Load the video or voice recording

The file with the talking in it. It is checked for an audio track straight away — that track is the trigger, so a silent clip cannot be ducked. Nothing is uploaded; the check runs in the tab.

2

Add the music and set the three controls

How far the music drops (Gentle, Standard or Heavy — each sets a threshold and a ratio: -20 dBFS at 4:1, -26 at 8:1, -34 at 16:1), how fast it reacts (5/150, 20/300 or 50/800 ms attack and release), and the music level before any ducking. If the music is shorter than the video the tool says so and offers to loop it.

3

Duck the music and download

FFmpeg sidechain-compresses the music against the voice, mixes the two, and copies the video stream through untouched. Output keeps your source container.

Popular task presets

Best for / not for

Best for

  • A talking-head video or a voice-over that needs a music bed which gets out of the way by itself.
  • Podcast video and tutorials, where the music has to sit under speech for half an hour without anyone riding a fader.
  • Any clip whose picture must not be touched: the video stream is copied, so only the audio is rewritten.

Not for

  • Silent clips. The voice is the trigger, so a clip with no audio track has nothing to duck against — the tool says so on load and points you at Add Music to Video.
  • A voice that never rises above the threshold the drop sets (-20, -26 or -34 dBFS). The run finishes, the file downloads, and the music never moves. Normalise the voice first.
  • Crossfaded loops. A short track is repeated with a hard join and cut dead at the end of the source; nothing fades in or out.
  • Ducking at moments you choose, or a second music cue later in the clip. The compressor reacts to the voice and nothing else.

When to duck music under a voice

  • A vlog or a tutorial where the bed has to stay audible in the gaps but never compete with the narration.
  • A voice-over for a slideshow or a product video, mixed once instead of drawn as a volume curve in an editor.
  • A long recording — an interview, a lecture, a podcast video — where automating the level is the only practical option.
  • A rough cut that has to be listenable before anyone opens a real DAW.

Supported formats & limits

Input containersMP4, MOV, WebM, MKV, AVI, M4V and the common audio containers — anything this FFmpeg build decodes. The file must have an audio track; a silent clip is rejected on load.
Input codecsH.264, H.265/HEVC, VP8, VP9, MPEG-4, MJPEG video with AAC, MP3, Opus, Vorbis, FLAC or PCM audio. Not AV1 — this build has no software AV1 decoder.
Output containerThe same container as the source: an MP4 in is an MP4 out, a MOV a MOV. Nothing is forced to MP4. An audio-only source comes back as an M4A.
Output codecVideo: copied bit-for-bit with -c:v copy, never re-encoded. Audio: re-encoded by the container default, or 192 kbps AAC when the source is audio-only.
What gets appliedsidechaincompress. The drop sets threshold and ratio together — 0.1 (-20 dBFS) at 4:1, 0.05 (-26 dBFS) at 8:1, 0.02 (-34 dBFS) at 16:1 — and the reaction sets attack/release to 5/150, 20/300 or 50/800 ms. The music level is a volume filter ahead of it.
The voiceNever compressed, and mixed back at its own level: the mix runs amix with normalize=0, so the finished audio is not 6 dB below the source. A mono voice spread to stereo loses about 3 dB on the way.
Max file sizeUp to ~2 GB (limited by browser memory) — and both files sit in memory at once, so the music counts too.
Max durationNo hard limit. The cost is re-encoding the audio; the picture is copied, so length matters far less than it does on a re-encoding tool.
CostFree for any use. No signup. No watermark.

Music Ducking vs. the usual alternatives

FeatureThis toolVEED (free)Kapwing (free)CapCut Online
Processing modelRuns locally in your browserUpload-based project editorUpload-based project editorUpload-based online editor
File limitsNo upload cap; practical limit is browser memoryPlan-specific upload limitsPlan-specific upload and export limitsFeature- and account-specific limits
Watermark on outputNo watermark addedFree exports include a VEED watermarkFree exports include a Kapwing watermarkStandard edits can be watermark-free; templates/assets may add branding
Signup / accountNo account for toolsWorkspace/account flowWorkspace/account flowCapCut account flow
Works offlineYes after cache, subject to browser supportNoNoNo
Best forPrivate one-step file operationsFull editor, templates, AI toolsCollaboration, templates, AI toolsSocial templates and timeline editing

Vendor plan limits were checked on April 29, 2026 and can change by region, account state, and export option. Verify critical limits on the vendor pricing/help page before relying on them.

Why use this music ducking tool

  • A real sidechain compressor, the same one an editor would reach for — not a fixed volume drop applied to the whole track.
  • The picture is copied, not re-encoded, so ducking the music costs no video quality and the run only takes as long as rewriting the audio.
  • The file is probed the moment you load it: no audio track, and it says so instead of handing you an unchanged file.
  • It prints the numbers it is about to use — threshold in dBFS, ratio, attack and release in milliseconds — so the result is reproducible in FFmpeg on the command line.
  • 100% private: the clip and the music are decoded by FFmpeg compiled to WebAssembly inside your browser, and neither is uploaded.

Task-focused FAQ

How much does the music actually drop?

Measured on a speech clip with the music at 40%: about 4 dB on Gentle, 9 dB on Standard, 14 dB on Heavy. The drop presets move the threshold as well as the ratio, because the ratio on its own is worth under 2 dB.

Why did nothing duck?

The voice never crossed the threshold. Gentle listens above -20 dBFS, Standard above -26, Heavy above -34, so a quietly recorded voice can run the whole clip without triggering anything. Normalise it first, or pick a heavier drop.

Is the video re-encoded?

No. The video stream is copied with -c:v copy, so the picture is bit-identical to the source and the output keeps the container you gave it.

Can I duck at specific moments instead?

Not here. The compressor follows the voice, and nothing else. Placing cues by hand is a timeline job.

Frequently Asked Questions

What does ducking mean?

Music that turns itself down while someone talks and comes back up in the gaps. A sidechain compressor does it: the music goes through the compressor, but the voice decides when the compressor clamps down. The voice is never compressed, and the mix runs with amix normalize=0 so it comes back at its own level rather than 6 dB below it. A mono voice spread across both channels loses about 3 dB on the way; a stereo one is unchanged.

Does my video need to already have sound?

Yes — the voice is the trigger. The tool probes the file on load and says so plainly if there is no audio stream. To put music over a silent clip at a fixed level, use Add Music to Video.

What if the music is shorter than the video?

You are told both durations before you run anything. By default the music loops to fill the video and is cut at the end — a hard join, not a crossfade. Turn looping off and the music stops where it ends while the voice carries on.

What exactly gets applied?

sidechaincompress. Each drop sets a threshold and a ratio together: 0.1 (-20 dBFS) at 4:1, 0.05 (-26 dBFS) at 8:1, 0.02 (-34 dBFS) at 16:1. Measured on a speech clip, that is about 4, 9 and 14 dB of ducking — moving the ratio alone, at one threshold, was worth 1.7 dB. Reaction speed sets attack/release to 5/150, 20/300 or 50/800 ms, and the music level is a volume filter ahead of the compressor.

Is the video re-encoded?

No. The video stream is copied with -c:v copy, so the picture is bit-identical and the run only costs as much as re-writing the audio. Audio-only sources come back as a 192 kbps M4A.

Why does it sound like it is pumping?

Usually a fast reaction with a heavy drop. Move the reaction to Natural or Slow so the music glides back instead of jumping, and step the drop down to Standard or Gentle.

Nothing seems to have ducked. Why?

Almost always a voice quieter than the threshold. Gentle only listens above -20 dBFS, so a voice recorded low never crosses it and the music holds its level for the whole clip — the run still finishes and still downloads. Normalise the voice first, or pick Heavy, which listens down to -34 dBFS.

Related Tools