πŸ”’ Your files stay on your device β€” core audio processing runs locally. Privacy details β†’
πŸŽ™οΈ
Free Β· Two local files Β· Envelope-based DSP

Auto Audio Ducker β€” Voiceover & Background Music

Mix narration over a music bed and automatically turn the background down while the voice is active. Set the threshold, duck depth, attack and release, preview the result and verify the actual reduction before exporting WAV or MP3.

Before-and-after illustration of Auto Audio Ducker: Voiceover and music becomes Ducked mix through Follow speech level.
Before → process → afterSee the signal story, then hear it in the workspace.

DYNAMICS MAP

Voice Ducking Envelope

Follow the level decision, not just the waveform. The important story is how the detector reacts, where gain starts moving and whether the result stays controlled without flattening the performance.

Before-and-after illustration of Auto Audio Ducker: Voiceover and music becomes Ducked mix through Follow speech level.
Voiceover and musicFollow speech levelDucked mix
A threshold-led gain path: peaks cross the decision line, the processor responds, and the output envelope becomes more deliberate.
Open the full signal breakdownBest source, processing logic, limits and three-pass listening check
  1. Start with the right material

    Podcast intros, explainers, tutorials, product videos and social clips where speech needs space over a continuous music bed.

  2. Understand the transformation

    The browser aligns decoded voice and music, measures a bounded voice envelope, builds an attack/release gain curve from the threshold and duck-depth settings, then mixes and renders the result locally.

  3. Verify the usable result

    A previewable WAV/MP3 mix plus a receipt for exact duration, active-voice share, maximum and average music reduction, output peak, clipping, sample rate and channels.

Music reduction

Where this workflow stops

This is envelope-based DSP, not AI dialogue understanding. Breaths, noise and quiet words can trigger differently; it does not replace a final mix, loudness pass or rights check.

THREE-PASS CHECK

Make the final decision by ear

  • Compare at matched loudness so β€œlouder” is not mistaken for β€œbetter”.
  • Inspect the loudest transient and the quietest useful detail.
  • Leave enough headroom for the next processing or delivery stage.
Inputs
One foreground voiceover or narration file and one background music file in formats your browser can decode.
Ducking controls
Voice threshold, maximum music reduction, attack and release timing, music level, and music-only intro/outro duration.
Output
A locally rendered voice-and-music mix, WAV or MP3 download and a receipt for measured ducking activity.

How to duck background music under voiceover online

  1. Select the clean voiceover or narration file that should control the mix.
  2. Select the background music file that should become quieter while the voice is active.
  3. Set the voice threshold and duck depth, then tune attack and release for natural transitions.
  4. Preview the mix, render it, inspect the measured receipt and download WAV or MP3.

Quick answer: Auto Audio Ducker uses the measured amplitude envelope of one local audio file to control the level of another. It is useful for voiceover and background music mixes where manual volume automation would be repetitive. It does not transcribe speech, understand words or use an AI model.

What threshold, duck depth, attack and release change

Threshold
Sets the voice-envelope level that activates ducking. A lower threshold reacts to quieter breaths and room noise; a higher threshold can miss soft words.
Duck depth
Sets the maximum background music reduction while the voice is active. Deeper reduction improves separation but can make the music pump or disappear.
Attack
Controls how quickly the music turns down after the voice passes the threshold. Short attack protects word beginnings; too short can sound abrupt.
Release
Controls how quickly the music returns after the voice falls below the threshold. A longer release sounds smoother; too long can hold the music down between phrases.

Start with a moderate duck depth, a fast but nonzero attack and a release long enough to bridge tiny gaps inside a sentence. Then listen to both quiet phrases and loud sections. The right settings depend on the recording, so a preset cannot guarantee a finished broadcast mix.

Measured ducking instead of a guessed result

The result receipt is designed to show what the render actually did, not just repeat the requested controls. It can report the resolved input durations and sample rate, the applied settings, how much of the timeline triggered ducking and the measured background gain reduction. Use the receipt alongside listening: a number can confirm processing activity, but it cannot decide whether the mix sounds natural.

Envelope-based, not AI: the detector responds to foreground signal level. Music, noise or breaths in the voiceover can also cross the threshold, while very quiet speech may not. Clean the controlling track first when false triggers are a problem.

Auto Audio Ducker versus Joiner, Compressor and BalanceDAW

Auto Audio Ducker
Mixes a foreground file and background file concurrently, then changes the music gain over time from the voiceover envelope.
Audio Joiner
Places files one after another. Use it for a sequence, not for narration playing over music.
Audio Compressor
Controls the dynamics of one signal from that signal itself. Use it to reduce its dynamic range rather than to make one file react to another.
BalanceDAW
Use the studio for multitrack timelines, detailed edits, manual automation, plugins or more than one voice-and-music pair.

Common voiceover and background music workflows

Limits to check before publishing

Privacy: the selected voiceover and background files are processed in this browser and are not uploaded to AudioWrench. Normal page assets and same-origin encoder code can still be requested as described in the privacy notice.

Auto Audio Ducker: quick answer and technical limits

Quick answer: A local voiceover and background-music mixer that lowers music when its envelope detector finds active narration, then restores it with adjustable timing.

Best for
Podcast intros, explainers, tutorials, product videos and social clips where speech needs space over a continuous music bed.
How it works
The browser aligns decoded voice and music, measures a bounded voice envelope, builds an attack/release gain curve from the threshold and duck-depth settings, then mixes and renders the result locally.
What you get
A previewable WAV/MP3 mix plus a receipt for exact duration, active-voice share, maximum and average music reduction, output peak, clipping, sample rate and channels.

Know before you use it: This is envelope-based DSP, not AI dialogue understanding. Breaths, noise and quiet words can trigger differently; it does not replace a final mix, loudness pass or rights check.

Privacy: Selected audio is processed in this browser and is not uploaded to AudioWrench. Normal page assets can still be requested as described in the privacy notice. Privacy details β†’

FAQ

How does automatic audio ducking work?
The tool follows the amplitude envelope of the voiceover. When that envelope crosses the chosen threshold, it lowers the background music toward the selected duck depth; attack controls how quickly the reduction begins and release controls how quickly the music returns. This is deterministic envelope-based DSP, not AI speech recognition.
What is this tool best used for?
It is designed for a voiceover or narration track mixed over background music, such as podcasts, video explainers, tutorials and spoken intros. It is not a full multitrack editor, and a noisy or highly dynamic voice track may need cleanup or manual automation for the best result.
How is Auto Audio Ducker different from Audio Joiner, Compressor and BalanceDAW?
Audio Joiner places files in sequence, while this tool mixes two files at the same time and changes the music level from the voice envelope. Audio Compressor changes one signal's own dynamics. BalanceDAW is the better choice for multitrack arrangement, detailed automation, edits and plugins.
Are my files uploaded, and can I export WAV or MP3?
The selected voiceover and music files are decoded, mixed and rendered on your device and are not uploaded to AudioWrench. WAV avoids another lossy encoding stage but is still a new browser render, while MP3 is a lossy local re-encode. Input codec support and practical file length depend on your browser and available device memory.

Related voice and mixing tools