Sound

Noise removal, and when to stop

The setting that sounds cleanest is usually the one that took the person out.

You found the noise reduction, dragged it up until the hiss went, and something else left with it. The room has gone. So has a bit of the voice, though you would struggle to say which bit.

Two habits save most recordings from this. Compare the cleaned version against the original rather than against silence, and stop one notch before it sounds clean.

The rest is why they work, and how to hear the damage while you can still take it back.

What does noise removal actually take away?

Noise removal comes in two broad shapes now. The older kind works from a profile: you give it a few seconds of the recording where nobody is speaking, it measures which pitches the noise occupies and how much of each, from the low rumble at the bottom to the airy hiss at the top, and subtracts that shape from the whole file. The newer, AI-driven kind needs no profile at all, having learned in advance what speech and noise tend to look like, and it separates them on its own. They are built very differently and they hit the same wall.

The trouble is that your voice is not somewhere else. The airy part of a hiss sits exactly where the s, f and th sounds live. A low rumble sits underneath the chest of a voice. Room tone, meaning the quiet continuous sound of a space with nobody in it, occupies very nearly everything. Take the noise away and you take pieces of the voice that were standing in the same spot.

That is not a flaw in one product. It is what separating one sound from another runs into, and everything in this category meets it, however the machine learning on top is described.

Why does the cleanest version sound wrong?

Because every voice you have ever enjoyed listening to was in a room. Rooms are how ears place people, and a voice with nothing around it reads as a voice that is not anywhere.

Nobody watching will tell you the audio was over-processed. They do not have the vocabulary for it and they were not listening for it. They will say it felt a bit odd, or they will simply stop watching, and you will go away and conclude that the topic was wrong.

What should you be listening to?

Not the noise. The noise is the thing you are trying to stop hearing, which makes it the worst possible thing to monitor.

Most people clean by soloing the gaps: play the pauses, check the hiss has gone, push further if it has not. That comparison has one direction to travel in and it ends at silence every time.

Do it the other way round. Play a sentence, switch the processing off, play the same sentence again, and ask one question. Does this still sound like the person? Not is it cleaner. Does the voice still have a body, do the consonants still finish, is there still air between the words. Judge it on the speech, with the picture running, because that is the only condition anybody else will hear it in.

And stop watching the waveform. It will tell you the noise has gone long before your ears agree, and it cannot show you the thing you are protecting.

How do you know you have gone too far?

Four tells. Any one of them means back it off.

1. The s sounds have picked up a lisp, a click or a faint whistle. 2. Breaths have vanished, so sentences begin out of nowhere. 3. The background swells up between words and ducks away when you speak, which is the processing breathing. 4. The voice sounds like it is behind glass.

Two gentle passes are kinder than one heavy one, because each makes a smaller guess.

And there is a reason to stop a notch early. You have been listening to that hiss for as long as you have been working on the file, and your tolerance for it is now zero. You are also the only person who will ever hear the before and the after. Everybody else gets one version with no reference. They will not miss a background you removed most of, and they will notice a voice that has been thinned.

What the automatic setting is optimising for

Automatic cleanup and the newer AI-driven tools are genuinely good, and they are good at one thing: finding where the noise lives and how much of it there is. That is measurement, and a machine measures faster and more accurately than a person with headphones and an evening.

What they cannot know is what the recording is for. Left alone they optimise for the absence of noise, because that is the only thing they can score. Quality and legibility get traded away to reach it, and the result passes every test the tool knows how to run.

Where the tools stop

At this studio those tools do the analysis and nothing else. What comes out, and how far, is decided by ear, one problem at a time, by somebody who knows what the video is for. That is not a philosophical position about machines. No measurement can tell you that this speaker's breath is part of how they teach, or that the room behind a cellist is worth more than the tidiness of the file.

The same judgement runs the other way, and that is the half worth taking from this. If a recording has a little steady background behind a voice that is close and clear, the honest move is to leave it alone. A quiet room behind a person is what a room sounds like, and it is doing work you would otherwise have to put back.

Where this decision actually gets made

Before you open the noise reduction again, play the raw file to somebody who was not in the room and ask what they notice. If they mention the background, clean it carefully and stop early. If they do not, you were about to spend an afternoon removing something only you could hear, at a cost only your viewers would pay.

Noise removal

Keep reading

Sound

Hum, fans and the fridge

Read the piece →
Sound

Why your video sounds echoey

Read the piece →
Sound

Can your audio be saved?

Read the piece →
Sound

How loud should music be under a voice?

Read the piece →
Where this gets done

This is a service, and the method is written down.

Everything above came out of doing the work rather than writing about it. If you want the method instead of the story, it runs in order on one page.

By
Dogu Arkan
· Updated
28 August 2026
See how BUBI does it
Start here

Start with a free audit

Tell us where your content is now. We will come back with what we would change and what result to expect.

A bare domain, a full URL or a channel link.
Received. We reply in writing, usually within three working days.
That did not send. Try once more, or write to us instead.

A person reads the channel and writes the audit by hand: a considered read typically takes three working days. That is the usual shape, not a promised turnaround. We use these details only to reply to you: no lists, no lurking.

What you will get

A fit snapshot: where your channel stands, and whether we are a match.

Two to three opportunities: specific, prioritised, yours to keep.

A recommended next step, even if that step is not us.

The audit is free and commits you to nothing: nobody follows up with a call you did not ask for.