How to Add Background Music to a Video: Keep Dialogue Clear (2026)
The export finishes, you press play, and the track you spent an hour choosing is buried under your own voice, or it steamrolls the voice completely. You can add background music to a video in under a minute. Getting the balance right is where most tutorials stop and where this one starts.
Music under speech is a mixing problem, not a styling choice. A bed that sits well below a speaking voice reads as produced; a bed sitting at the same level reads as broken. Below are the two methods that actually work, the level targets to aim for, and the licensing rules that keep an upload from being muted.
You do not need expensive software. You need to recognise what “too loud” sounds like and know the one control that fixes it.
The short answer: drop the music, then duck it
Background music under speech is a level problem before it is anything else. The bed should sit roughly 15 to 20 decibels below the voice, and it should fall further whenever someone talks. That automatic drop is called ducking.
Most editors can do it, and so can FFmpeg. If you remember one number, remember this: the music should be quiet enough that you stop noticing the track and start noticing the words. If you have to strain to hear the speaker, the mix is wrong.
Both routes below finish in the same place. Pick the one that matches the software you already have.
What you need before you start
You need three things: your video with its existing voice audio, a music file you are allowed to use, and a tool that can mix two audio streams.
Either tool lets you add background music to a video, but the workflow differs. An editor with an audio mixer (DaVinci Resolve has a free version, and so do Shotcut and Kdenlive) gives you a timeline. FFmpeg, the command-line tool for Windows, macOS and Linux, gives you a repeatable command. Your editor’s own help pages explain where its ducking control lives, because menus differ between versions.
The music file matters more than the software. A track with a clear, uncluttered midrange, such as soft piano or a simple beat with no vocals, sits under speech far more easily than a dense pop mix. Instrumental versions beat full songs almost every time.
Method 1: add background music to a video and duck it in an editor
This is the route for most people, because you can see the waveforms and hear the result as you work.
- Import your video, then drop the music file onto a second audio track beneath the dialogue.
- Pull the music clip gain down until its peaks sit well under the voice. Start near minus 18 dB and adjust by ear.
- Turn on ducking. In editors that support it, this is a sidechain or auto-duck setting on the music track, triggered by the dialogue track. In editors without it, lower the music by hand during every spoken section.
- Fade the start and end of the music so it does not begin or stop with a click. One to two seconds usually sounds natural.
- Play the whole thing on a phone speaker, then on headphones. Phone speakers lose bass and expose a muddy mix; headphones expose a bed that is too loud.
- Export with the video stream copied or lightly compressed, and the audio re-encoded as AAC at 192 kbps or higher.
If your editor has no ducking feature at all, do not skip it. Volume keyframes around each sentence are tedious, but they still beat a static music level.
Method 2: duck music under a voice with FFmpeg
FFmpeg ships a filter called sidechaincompress. It listens to the voice track and compresses the music whenever the voice is present, which is exactly what ducking means.
One command adds the bed, ducks it, and leaves the picture untouched:
ffmpeg -i video.mp4 -i music.m4a -filter_complex "[0:a]volume=1.0[voice];[1:a]volume=0.3[music];[music][voice]sidechaincompress=threshold=0.03:ratio=8:attack=20:release=400[ducked];[voice][ducked]amix=inputs=2:duration=first:dropout_transition=0,afade=t=in:st=0:d=1.5[aout]" -map 0:v -map "[aout]" -c:v copy -c:a aac -b:a 192k output.mp4
Treat those numbers as a starting point, not a rule. Threshold decides how loud the voice must be before the music ducks; ratio decides how hard it ducks; attack and release control how fast it reacts and recovers. A short release makes the music pump in and out between words.
The official FFmpeg filter documentation lists every sidechaincompress option, including makeup gain and the level_sc input control. If the command feels like too much at once, mix at a fixed level first, then add ducking once the basic balance sounds right.
What the tools in this space actually do, and where they stop
Search for adding music to a video and two kinds of tool appear. The first is real editing software with an audio mixer and ducking. It does what you tell it, and nothing more.
The second is the online “add music” page: upload, pick a track from a built-in library, download. These are genuinely convenient for a slideshow or a short social clip, and limited everywhere else. They often cap file size or length, they re-encode your video, and the music they supply comes with a licence that governs how you may publish the result.
Neither kind can usually fix a voice recording that is already noisy, and neither will license a chart song for you. Ducking is arithmetic on levels; it cannot rescue a mix that began with poor source audio. If room tone is the problem, clean that up first, and our guide to removing background noise from video covers that pass.
Loudness targets are worth knowing too. Most platforms normalise playback toward roughly minus 14 LUFS, and a mix sent in much hotter is simply turned down, taking the dialogue with it. The reference most tools follow is the EBU R128 loudness standard. Set the level of the final mix, not just the music.
Rights and responsible use
Adding a track to your timeline changes nothing about who owns it. The music belongs to its composer, performers and rights holders, and placing it under your video transfers none of that to you.
Personal reference is one thing, publishing is another. A private edit you never upload sits in a different category from a video you post publicly, monetise, or hand to a client. The moment a track reaches an audience, the licence matters.
Use music you are actually licensed to use. That means an official library such as the audio library inside the platform you upload to, a paid royalty-free licence that names your use, or music you commissioned or wrote yourself. Owning a copy of a song, or paying for a streaming subscription, is not a licence to reuse it.
Uploaded content is scanned automatically. On YouTube, for example, the Content ID system matches audio against a reference database, and the rules are set out in YouTube’s Terms of Service. They apply to the audio under your video exactly as they apply to the video itself.
Troubleshooting
The music drowns out the dialogue
The bed is too high, or ducking is not triggering. Drop the music another 6 dB, then confirm the trigger is the dialogue track and not the final mix.
The music pumps in and out between words
The release time is too short, so the bed snaps back up in every pause. Lengthen the release to around 300 to 500 ms and soften the ratio.
It sounds balanced on a laptop but wrong on a phone
Laptop speakers flatter the midrange, while tiny phone speakers lose bass and make speech harder to follow. Mix for the worst speaker you expect, and check both.
The music starts or ends with a click
A hard cut on a waveform produces a click. Add a fade of at least one second to the head and tail of the music.
The upload is muted or claimed
The track matched a rights holder’s catalogue, which is a licensing problem rather than an editing one. Replace the music with something you are licensed to use.
The music sounds muffled or the pitch is off
Your music file’s sample rate probably disagrees with the project. Resample the music to the project rate, usually 48 kHz, before mixing rather than after.
Frequently asked questions
Can I add background music to a video for free?
Yes. Free editors such as DaVinci Resolve, Shotcut and Kdenlive mix two audio tracks, and FFmpeg is free and open source. The software is rarely the cost; a licence for the music usually is.
How loud should background music be under a voice?
Start 15 to 20 decibels below the dialogue and duck it further while the person speaks. If you notice the track more than the words, it is too loud.
Do I need a paid editor to duck music automatically?
No. Several free editors include ducking or sidechain tools, and the FFmpeg filter ships with the free build. Paid software mainly saves time, not capability.
Will adding music get my video claimed on YouTube?
Only if the track matches a rights holder’s catalogue. Audio from the platform’s own library, or a properly licensed track, is much less likely to be claimed for the audio itself.
Can I use a song I bought on iTunes or stream on Spotify?
No. A purchased copy or a streaming subscription gives you listening rights, not the right to place that recording under your own video and publish it.
What is the fastest way to add music to a video on my phone?
A phone editor with a built-in music library is fastest, because the app supplies both the mixer and a licensed track. Expect fewer level controls than a desktop editor.
How do I fade music in and out?
Apply a one to two second fade at the start and end of the clip. In FFmpeg that is the afade filter, as shown in the command above.
Which method should you use?
If you want to watch the waveform and hear changes as you make them, use an editor. It is the safer choice for a one-off video and the easiest place to learn what a good balance sounds like.
If you are processing many clips, working on a server, or you simply prefer a keyboard to a timeline, the FFmpeg sidechaincompress route is faster to repeat and script.
Whichever route you choose, the way you add background music to a video matters less than the levels you set. Get the voice right, place the bed 15 to 20 decibels under it, duck while the person talks, fade the edges, then check the whole thing on a phone speaker before you publish.
When the balance is settled, you can normalise the final levels so the video sits at a consistent loudness beside everything else on the platform. If you want more command-line recipes for the same workflow, our everyday FFmpeg commands collection has them.