tools / audio / voice over music

Voice over music

Narration over music with auto-ducking, the podcast mix without the DAW.
Runs locally in your browser. Enable JavaScript to use the interactive tool. The practical details remain available below.

What Voice over music actually does

Ducking means the music moves: full level in the gaps, down under the voice, back up afterwards. A control that simply makes the music quiet for the whole piece is not ducking, and it is a common substitute because it is easier to implement and hard to tell apart from a description. This page keeps the two ideas separate. The resting level of the music is one control, and how far the music drops under the voice is another, given in decibels. The depth is reached by measuring how loud the voice actually is and solving the compressor's threshold against that reading, because a fixed threshold produces a different depth on every recording.

How to use it

  • Choose the voice first and the music second. Both appear on one timeline at one shared scale.
  • Set how far the music should drop while the voice speaks, and its resting level for the gaps.
  • Mix, then read the panel showing the measured voice peak and the threshold that was solved from it.

Useful for

  • Mix a podcast intro where the bed drops under the host and comes back between sentences.
  • Put a music bed under a recorded voiceover for a video without editing by hand.
  • Produce an audio ad where the bed has to stay audible but never compete with the read.

Limits worth knowing

  • Exactly two files are used, voice then music. A third is refused rather than silently ignored.
  • The depth is solved from the peak level of the voice as a whole. A recording whose level varies widely will duck less in its quiet passages than in its loud ones, which is usually what is wanted but is not a fixed number.
  • When the voice cannot be measured the page falls back to a documented default and says on the result panel that it did.

Questions people ask

Are the files uploaded?

No. The measurement pass, the mix and the read-back all happen in this tab. The audio engine is fetched once and cached.

Why is the music no longer quiet throughout?

Because that was not ducking. The music now sits at its own level and drops only while the voice is present, by the depth set on the page.

What is the threshold figure on the result panel?

It is the level the compressor treats as the point where the voice starts pushing the music down, solved from the measured peak of the voice so that the requested depth is reached.