Audio Mixing Fundamentals for Video
Learn the core principles of audio mixing for video: balancing dialogue, music, and SFX using levels, panning, compression, and proper metering standards.
Introduction to Audio Mixing for Video
Audio mixing is the process of combining all audio elements in a video project — dialogue, music, sound effects, ambiance — into a cohesive, balanced, and emotionally effective soundtrack. A well-mixed video sounds effortless and natural; a poorly mixed one is immediately jarring, whether the music drowns out the dialogue, sound effects are too loud, or the overall levels are too quiet or too hot.
This tutorial introduces the fundamental concepts and practical techniques of audio mixing specifically for video content, from short-form YouTube videos to documentary films.
The Three Pillars of a Video Mix
Every video has three primary audio layers:
1. Dialogue (D)
Dialogue is always the highest priority. Viewers can tolerate imperfect music and thin sound design, but if they cannot understand what someone is saying, they will stop watching. Dialogue should be clear, intelligible, and consistent in level throughout the video.
2. Music (M)
Music sets the emotional tone, drives pacing, and fills acoustic space. In most video contexts, music lives below dialogue and should support it, not compete with it.
3. Sound Effects and Ambiance (SFX/FX)
Sound effects create realism (footsteps, doors, transitions) and atmosphere (ambiance, room tone). They reinforce the visual and emotional impact of scenes without drawing attention to themselves in most cases.
Understanding Audio Levels
Decibels (dB)
Audio levels are measured in decibels. Key reference points:
- 0 dBFS — the absolute maximum a digital system can handle. Any signal that hits 0 dBFS is clipping (distorting). Avoid this at all costs.
- -6 dBFS — a safe ceiling for individual tracks. Leave headroom above each element.
- -23 LUFS — the international broadcast standard for integrated loudness (EBU R128 / ITU-R BS.1770). Most YouTube, streaming, and broadcast platforms normalize to around -14 to -16 LUFS.
Target Levels by Element
A practical starting point for a narrative or documentary video mix:
| Element | Target Level |
|---|---|
| Dialogue (peak) | -6 to -3 dBFS |
| Music (under dialogue) | -18 to -24 dBFS |
| Music (solo, no dialogue) | -12 to -9 dBFS |
| Sound effects (design) | -12 to -6 dBFS |
| Ambiance / room tone | -30 to -24 dBFS |
Setting Up Your Mix
Organize Your Tracks
Before mixing, organize your timeline:
- Create dedicated audio tracks for each category: Dialogue, Music, SFX, Ambiance.
- Color-code tracks for quick visual identification.
- Route each category to a submix bus (an intermediate mixing channel). This lets you control all dialogue at once with a single fader, for example.
Set Gain Staging
Gain staging means setting appropriate levels at every stage of the signal chain:
- Clip gain — set individual clip volumes so that the loudest moments peak around -12 to -6 dBFS before any effects are applied.
- Track fader — use the fader to fine-tune relative balance between elements.
- Bus fader — control the overall level of each category (Dialogue bus, Music bus).
- Master fader — the final output level. Aim for a peak around -3 dBFS and an integrated LUFS around -14 to -16 for most online delivery.
Mixing Dialogue
Consistency is Key
Dialogue recorded across multiple days, locations, or microphones will have inconsistent levels and tone. Before any other processing, go through all dialogue clips and normalize or manually adjust clip gain so all clips have similar volume.
Essential Dialogue Processing Chain
Apply this signal chain to dialogue tracks (in order):
- High-Pass Filter (80–100 Hz) — removes rumble and low-frequency noise.
- EQ — address any tonal problems specific to the recording.
- Compressor — evens out the dynamic range. A ratio of 2:1 to 4:1, attack of 10–20ms, release of 100–200ms, and threshold set to catch the louder peaks is a safe starting point.
- De-esser — if sibilance is an issue, place a de-esser after compression.
- Limiter — a safety limiter at -3 dBFS to prevent any peaks from clipping.
Mixing Music
Choose the Right Music
Before mixing, choose music that works for the emotional arc of the scene. High-energy music with heavy beats will require more volume reduction under dialogue than soft, understated underscore.
Automating Music Volume
Static volume levels rarely work well for music. Use automation to smoothly duck music whenever dialogue begins and raise it back when dialogue ends:
- Enable track automation recording in your editor (write mode).
- Play the timeline and manually ride the music fader down when dialogue starts, up when it ends.
- Review the automation curve and smooth out abrupt jumps.
- Alternatively, use a sidechain compressor or a ducking plugin to automate the dip automatically when the dialogue track has signal.
Music Editing for Video
Match music edits to scene cuts wherever possible. Cut or crossfade music at a beat or phrase boundary rather than mid-phrase to avoid an abrupt ending.
Mixing Sound Effects
Layering SFX
Complex sound design often requires layering multiple elements for a single effect — for example, a punch sound might combine a low thud, a mid smack, and a high-frequency impact spike. Mix each layer individually:
- Give lower elements slightly more volume and less high-end.
- Give the "body" layer the most presence.
- High-frequency impact layers should be subtle to add definition without harshness.
Room Tone and Ambiance
Every location has a unique ambient sound. Match room tone under all dialogue scenes — a moment of silence in an interview should not be completely dead; it should have the same background hum as the surrounding audio. Use room tone clips (recorded on set) to fill any gaps.
Metering and Delivery Standards
True Peak vs. LUFS
Modern audio uses two meters:
- True Peak — the actual peak level of the signal. Never exceed -1 dBTP for deliverables.
- Integrated LUFS — the average perceived loudness over the entire piece. Target -14 LUFS for YouTube, -16 LUFS for streaming, and -23 LUFS for broadcast.
Check your delivery platform's specific loudness specifications before exporting.
Exporting Audio
- Export a stereo mix for most online video platforms.
- Use 48 kHz / 24-bit WAV for broadcast and professional delivery.
- Consider exporting a separate Music and Effects (M&E) stem for international versioning.
Quick Reference: Common Mixing Mistakes
- Music too loud — the single most common mistake. If you can't easily understand speech, the music is too loud.
- No automation — static levels sound mechanical. Always use volume automation for natural transitions.
- Clipping on the master — always leave headroom. Check your master output meter throughout the mix.
- Inconsistent dialogue levels — address clip gain before riding faders.
- Ignoring room tone — silence in video is obvious and unnatural. Always fill silence with appropriate ambiance.
Conclusion
Audio mixing for video does not require years of music production experience — it requires a clear understanding of priorities (dialogue first), proper level management, and consistent application of a few core tools: EQ, compression, and volume automation. Start with a well-organized timeline, set sensible gain staging, and work from the most important element outward. With practice, a clean, professional mix becomes a reliable part of your editing workflow.
Try it yourself — free in your browser
No upload, no signup, no watermark — these tools run on FFmpeg WebAssembly locally.