Sound

Mixing a lesson so people can follow it

Satisfy either competing goal fully and you have failed the other, so the work is deciding which one is the lesson.

You film yourself playing something, or cooking something, or taking something apart, and then you talk over it. Two sounds, one video, and every version you try makes one of them worse.

That is not a failure of your mixing. It is the shape of the problem. Decide which of your two sources is the lesson, protect that one completely, and accept that the other is now in service of it. A lesson that sounds gorgeous and cannot be followed has failed, and so has one that is perfectly clear over an instrument you have flattened into cardboard.

Which two goals are fighting?

The first is that the thing being demonstrated should sound the way it deserves to sound. A piano should have weight, decay and some room around it. A pan should sizzle like a pan. Whatever you are showing people, you showed it because it is worth showing.

The second is that the teaching on top has to stay legible. Every word, on a phone, at half volume, to somebody who does not yet know the vocabulary and cannot fill in a word they missed.

Both are legitimate, and people who are good at one are usually deaf to the other. Musicians mix the instrument. Teachers mix the words. A video needs somebody holding both at once, which is uncomfortable, and being uncomfortable is what the job consists of.

Why can you not satisfy both?

Two mechanisms, and they stack.

The first is masking, which is the plain fact that when two sounds occupy the same pitches at the same moment, the louder one hides the quieter. Speech lives in the middle of the range. So does the body of nearly every instrument, and so does the sound of a hot pan. That collision is closer to physics than to taste: the exact rules of what hides what are more tangled than a straight sum, but no amount of care over your headphones talks two sounds out of the same part of the range.

The second is attention, and it is the one that catches people who have already solved the first. Most people hold a single stream in the foreground at a time, and a viewer who is trying to learn is not going to split their attention to rescue your sentence. Make the instrument beautiful enough and they will listen to the instrument, and your sentence will go past intact, audible and unheard. That is the failure which looks like success: every word technically clear, nothing learnt. You cannot catch it in your own edit, because you already know what you said. You catch it by watching one other person watch it.

Which of your two sources is the lesson?

Ask what a viewer would complain about if it were missing.

If they came to learn the piece, the playing is the lesson and the talking is signposting. If they came to learn a technique, the words are the lesson and the playing is the illustration of it. Most people filming lessons have never made that choice out loud, and a video that has not made it drifts between the two, which is why it is tiring to watch and impossible to mix.

It can also change inside a single video, and usually does. Two minutes of teaching and then thirty seconds of demonstration are two different balances, not one compromise held across both.

What does protecting the lesson cost?

Something real, which is why the decision has to be made rather than avoided.

During the teaching passages, the demonstration steps back. It gives up a little of its low end and a little of its width so the middle of the range is clear for the voice. It will not sound as good on its own, and it is not being heard on its own.

During the demonstration, nobody talks, and everything comes back: weight, width, the room it was played in. The join between those two states is where home mixes give themselves away, because one setting has been applied to the whole video, so either the demonstration is thin or the teaching is buried.

The largest single improvement available here is not a mix decision at all. Stop talking over the best part. Give the demonstration thirty seconds of its own with nothing on top, and the video improves more than it would from a week of balancing. It costs an edit and nothing else.

There is one recording decision behind all of this, and it is worth knowing before your next shoot. If the voice and the instrument arrive on the same recording, their balance is close to fixed. The newer separation tools can prise the two apart to a degree, but imperfectly and with artefacts, so two separate recordings remain the only way to keep the decision genuinely open for as long as the project exists.

Does this apply outside music?

It applies to anything with a sound worth hearing underneath somebody explaining it.

A cook talking over a pan. A mechanic over a running engine. A potter over a wheel. A coach over a room of people. The question is the same every time and the answer is not always the voice, which is the part people get wrong. In a video about how a particular fault sounds, the engine is the lesson and the words are labels on it, and mixing it the usual way produces a video where a diagnostic sound has been politely turned down until it can no longer be diagnosed.

Where the balance stops being a setting

The way out is two settings rather than one. The demonstration keeps its depth where nobody is talking, and gives some of that depth up where the teaching has to carry. One setting can only serve one of the two goals honestly, and a setting that serves both equally serves neither.

Which is the decision this hands you. Watch your last lesson back with a pen, and mark where you are teaching and where you are demonstrating. If those two are on top of each other for most of the runtime, no mix was ever going to rescue it, and what the video actually needs is thirty seconds of quiet in the right place.

Audio mixing

Keep reading

Sound

How loud should music be under a voice?

Read the piece →
Sound

Recording an instrument and a voice together

Read the piece →
Sound

Do you need more than one microphone?

Read the piece →
Sound

Why your audio is quiet on YouTube

Read the piece →
Where this gets done

This is a service, and the method is written down.

Everything above came out of doing the work rather than writing about it. If you want the method instead of the story, it runs in order on one page.

By
Dogu Arkan
· Updated
28 August 2026
See how BUBI does it
Start here

Start with a free audit

Tell us where your content is now. We will come back with what we would change and what result to expect.

A bare domain, a full URL or a channel link.
Received. We reply in writing, usually within three working days.
That did not send. Try once more, or write to us instead.

A person reads the channel and writes the audit by hand: a considered read typically takes three working days. That is the usual shape, not a promised turnaround. We use these details only to reply to you: no lists, no lurking.

What you will get

A fit snapshot: where your channel stands, and whether we are a match.

Two to three opportunities: specific, prioritised, yours to keep.

A recommended next step, even if that step is not us.

The audit is free and commits you to nothing: nobody follows up with a call you did not ask for.