How loud should music be under a voice?
There is no correct number, but the method is quick, free, and survives changing the track.
You have written something, recorded it, and put music underneath, and now you cannot tell whether the music is too loud. You turn it down. You listen again. You turn it up. An hour later you are back where you started and you have stopped being able to hear either of them.
The method is quicker than the hour you just spent. Set the voice first and then do not touch it again. Bring the music up from nothing until the moment you notice it, then take it back down until you stop noticing it. That is the level.
Then check it on a phone speaker, because your phone will disagree with your headphones and one of them is telling the truth about your audience.
Is there a correct level?
No. It depends on how dense the track is, where its energy sits compared with your voice, how close your voice was recorded, and where the video will be watched. Change any one of those four and the number moves, so a post that hands you a figure has quietly assumed all four.
The wider craft does carry a rough convention, music set a few decibels under the voice, and it is a reasonable place to leave the fader before you start listening. It is a starting point and nothing more. Under a dense track it is nowhere near far enough.
One number nearby is handled differently, and it is worth not confusing with this one. How loud the finished video arrives is set by each platform's own normalisation rather than by taste, a target that is real, and one that some destinations publish (Spotify names about -14 LUFS) while others such as YouTube leave you to measure, and that is a different question with its own post. This page is only about the balance between the music and the voice, which has no such number at all.
How do you set it, then?
Voice first, and then leave it. Most of the hour people lose to this is spent moving both things at once, so every adjustment changes the thing being measured. Fix one and the problem halves.
Then bring the music up from silence rather than down from loud. Creep it up until you become aware of it as music, mark that point, and come back down until it recedes and you are aware only of a voice with something underneath. Coming down from loud works less well, because your ears have accepted the music already and you are lowering it against a reference that has moved.
The test uses attention rather than legibility because you cannot judge your own legibility. You know what the words say, you cannot un-know the script, and you will always follow it. What you can still feel is the moment music stops being a background and becomes an object.
Two free checks finish it. Play the video from another room while you do something else and see whether you can still follow the sentences from the kitchen. Then show it once to somebody who has not read the script.
Why do headphones, a phone and a car disagree?
A phone speaker has almost no low end, so the bass of a track disappears and the middle of the range, where speech lives, arrives at full strength. A balance that felt gentle on headphones turns into a fight there.
A car has road noise and plenty of bass, so it buries quiet detail and promotes anything with weight. Headphones show you everything, including what nobody else will ever hear, which makes them the best place to find faults and the worst place to take the final decision.
If you can only check one, check the phone, at the volume you would actually use, somewhere with a bit of noise in the room.
Why does the track you chose matter more than the fader?
Because most of what decides this was chosen before you touched a level.
A track with a vocal in it competes for the same attention as your voice and will fight at every setting. A track with a busy middle does the same thing more quietly, since that is precisely where speech lives. A sparse track, with its energy at the bottom and the top, leaves a gap in the middle that a voice sits in comfortably, and it can be noticeably louder without ever getting in the way.
A track that changes every eight bars will also pull attention every eight bars. Swapping the track is very often the largest single improvement available, and if your library is already paid for it costs nothing.
What is ducking, and when does it sound mechanical?
Ducking means the music steps back automatically whenever the voice arrives and comes back up when the voice stops. Some editors do that with a compressor keyed off the voice and some with a keyframed volume that just follows it, so the same word covers two different mechanisms. Set well, nobody notices it. Set badly, it is the most conspicuous thing in the video.
It goes wrong when it is too deep or too fast. The music lurches up in every gap between sentences, drawing attention to the moments you wanted invisible, and it swells after every breath. Shallower and slower usually settles it.
The honest rule of thumb: for continuous narration, one well chosen level set by hand beats ducking, because there are no gaps long enough for the music to earn coming back. Ducking pays for itself where speech is intermittent over a long bed, since there the music has to be present between the lines and gone during them.
When the music should not be there at all
Some videos are better with nothing underneath. A dense lesson, with instruction in most sentences, gets very little from a bed and pays for it in the one currency it cannot spare.
Music that runs for the whole video also stops being heard before long, so it is doing no work while still costing legibility. Bringing it in and out, arriving at a change of subject and leaving when the detail starts, is worth more than any amount of care over its level.
Which is the question to answer before you touch the fader again. Name the job the music does in this passage: holding a gap, marking a change of subject, carrying a demonstration. If you can name it, the level stops being a mystery. If you cannot, the track should probably not be under that passage at all.
Keep reading
This is a service, and the method is written down.
Everything above came out of doing the work rather than writing about it. If you want the method instead of the story, it runs in order on one page.
Start with a free audit
Tell us where your content is now. We will come back with what we would change and what result to expect.
A person reads the channel and writes the audit by hand: a considered read typically takes three working days. That is the usual shape, not a promised turnaround. We use these details only to reply to you: no lists, no lurking.
What you will get
A fit snapshot: where your channel stands, and whether we are a match.
Two to three opportunities: specific, prioritised, yours to keep.
A recommended next step, even if that step is not us.
The audit is free and commits you to nothing: nobody follows up with a call you did not ask for.