Guide

How to Fix Bad Audio: What Comes Out, and What Never Will

A two-minute listening test that tells you whether a bad recording can be rescued, which noises lift off a voice, and which ones are baked into it for good.

Updated August 6, 2026

There is only one question in front of you, and it is not how to fix the recording. It is whether the recording can be fixed at all. The guest has gone home, the event packed up hours ago, the call happened once, and the answer decides whether you publish this as it stands, spend tonight on it, or write the session off.

Some of what you can hear is a separate layer sitting under the voice, put there by a fan or a fridge or a badly placed cable. Layers lift off. The rest is the voice itself, bent out of shape on the way in by the room or by the recorder, with no clean version underneath to recover. Nearly everything people call bad audio is one or the other, and two minutes with headphones will tell you which one you have.

Play the parts where nobody is talking

Use headphones or earbuds. Laptop speakers hide most of what you are trying to identify and phone speakers hide the rest.

Find a gap. The second and a half between two sentences, the moment before somebody answers a question, the run-up before the first word. Play it and listen to what is left when the voice is not there.

Whatever you hear in that gap has been running the whole time. You could not hear it under the talking, but it never stopped. A sound that holds steady like that can be measured while nobody speaks and then subtracted from everywhere else, which is all noise removal really is. The fan comes off the file and the voice stays where it was.

Now play a full sentence and listen for the same sound.

If it arrives with the first word and stops after the last, it is not underneath the voice. It is being made by the voice: the room throwing it back, or the recorder mangling it on the way in. Nothing can lift out a sound whose only source is the sound you want to keep.

That is the entire test. Constant in the gaps, removable. Only present when someone is talking, permanent.

What you are hearing, and whether it comes out

Most recordings have more than one thing wrong with them.

These come off.

  • A hum, buzz, or whirr under everything. Air conditioning, a fridge two rooms away, a laptop fan, a desk full of power supplies. Steady, boring, and precisely what noise removal exists for.
  • A hiss like blank tape. Same treatment, same result. Worth knowing where it came from, though. Hiss usually means the microphone was turned up hard, which usually means the voice was too far away from it. The hiss comes off. The distance does not.
  • Clicks, pops, and mouth noise. The sticky little sounds between words get a pass of their own.
  • One person much louder than the other. The most common complaint about any two-person recording and the least dramatic thing on the list to fix.
  • The whole thing too quiet. Usually, though quiet and distant are easy to confuse and only one of them comes back.

These do not.

  • A single bang. A door, a cough, a phone buzzing against the table. Half a second of noise is sitting on top of half a second of speech, and no filter takes one without taking the other. What that needs is a cut, which is an edit rather than a cleanup.
  • A hollow, boomy, bathroom quality. That is the room, and this cleanup does not take a room back out of a recording.
  • Crackle or fizz on the loud words only, clean everywhere else. The recorder ran out of ceiling. The loudest moments hit the top of what it could store and were flattened off there, which recordists call clipping, and the flattened part was never written to the file at all. Turning it down afterwards gives you a quieter crackle.
  • Music, a television, or a second conversation behind the voice. All of those move the way a voice moves, so there is no steady layer to measure and nothing to subtract. Pulling a finished mix apart into separate tracks is a different job, and not one Chat Octopus does.
  • Words that are missing, stuttered, or robotic. A dropped connection does not leave a damaged sentence. It leaves no sentence. If the speech in a file is not intelligible, Chat Octopus stops and asks you what you want to do rather than guessing at words that were never recorded.

"Too quiet" and "too far away" sound identical until you turn it up

These two get mistaken for each other constantly, and the difference between them is the difference between a one-line fix and a lost afternoon.

A quiet recording is a clean voice stored at a low level. Everything that happened in the room is on the file in the right proportions. The whole thing is simply small, and making it bigger is arithmetic.

A distant recording is a different balance. Stepping away from a microphone does not just make the voice quieter. It makes the voice quieter while the room carries on at exactly the volume it always was. What you have is a voice with more room mixed into it, and volume does not change a mixture. Turn it up and you get a louder recording of somebody in the next office.

Ten seconds settles it. Raise your own player until the voice sits at a normal listening level, then listen to what came up alongside it.

  • Voice normal, background still quiet. You had a quiet recording. Nothing to worry about.
  • Voice normal, hiss or room noise now obvious. The quiet layer under the voice rose by exactly as much as the voice did, because it always does. Removable, though how much gets removed is capped by how natural you want the voice to sound afterwards.
  • Voice normal, still sounds like it is coming from across the room. That is distance, and this is as good as it is going to get.

Why an echoey room never comes out

Reverb, the reason a kitchen sounds different from a car, is your own voice arriving more than once. It reaches the microphone directly, then again a few thousandths of a second later off a wall, then off a window, a tabletop, the ceiling, from every direction, hundreds of copies each fainter than the last. Hard flat surfaces make many strong copies. Curtains, sofas and full bookshelves swallow most of them.

Every other problem here is one sound laid over another. Reverb is one sound plus itself. There is no separate layer to strip away, because every copy is made of the same material as the original, and subtracting it would mean subtracting the voice. Anything that goes after reverb has to work on the voice itself, trading one problem for another, and a cleanup here does not make that trade.

That is why boomy recordings stay boomy. It is also why the effect gets worse the further a speaker sits from the microphone: the direct sound falls away with distance and the reflections do not.

If yours is echoey, a cleanup will still take out the fan and even up the levels. It will not make the room smaller. Judge the recording on whether the echo is tolerable, not on whether it will be gone, because it will still be there.

When the audio cannot be saved, the words usually can

A recording is two things at once. It is a sound file, and it is everything that was said. Words survive conditions that audio does not, and speech that is unpleasant to sit through is very often still perfectly legible.

So before you write the session off:

  • Turn it into text. A transcript of a badly recorded interview is an entirely good interview once it is on the page, with the quotes, the numbers and the timings intact. Running the cleanup first still pays off even when it fails to save the recording, because a steadier, quieter file comes back as a transcript with fewer mistakes to chase down.
  • Put the words on screen. If the audio belongs to a video, captions carry a clip whose sound is rough, and most of the people watching it on a phone in public were never going to hear it anyway.
  • Replace the voice, keep the pictures. This only works when the voice is yours to record again, but recording the narration again over the original pictures sidesteps the problem instead of fighting it.

What actually happens to a recording that can be fixed

There is a rough order to it, and the first step is not a filter.

The words get written down first. Before a single filter runs, the recording is transcribed down to individual word timings, so there is a record of what was said and nothing can quietly go missing later on.

Then noise removal, deliberately gentle. The rule it works under is that audible noise removal has already gone too far. A recording with a bit of room left in it still sounds like a person; pushed harder you get pumping, where the background swells up between words and ducks away underneath them, and nobody who has once noticed that can stop noticing it.

Then dead air. Any pause longer than about seven tenths of a second comes down to roughly three tenths. Breaths stay, beats stay, and the long silence while somebody remembers what they were going to say does not.

Then levels, in two passes. Voices are evened out against each other before anything else, so a guest recorded at half the host's volume is fixed against the host rather than against the file as a whole. Only then does the recording get its overall target, about -16 LUFS integrated. Integrated means averaged across the whole file rather than measured at any one moment, and LUFS is the scale every podcast and streaming app normalizes on, which is what stops your episode arriving quieter than whatever played before it. Minus sixteen is the spoken-word target; some platforms settle a couple of points louder. Peaks are held under -1 dBTP in the same pass, the margin that keeps the single loudest instant in the file from crackling on somebody else's speakers.

Last, a light EQ pass, which means turning particular parts of the sound's range up or down rather than the whole thing at once. The rumble below 80 Hz rolls off, the range where desk bumps and passing traffic live and a voice has almost nothing to contribute. Harsh S sounds get softened. A dull voice gets a small lift between 2 and 5 kHz, the band that makes speech sound close rather than muffled. All of it stays small on purpose. The reason so much cleaned-up audio sounds cleaned up is EQ applied with a heavy hand.

One rule overrides all of that. If the processing starts making the voice sound artificial or robotic, it backs off and leaves noise in instead. A slightly noisy recording that still sounds like a human being beats a silent one that sounds synthetic, every time.

What comes back matches what went in: an MP3 at 192 kbps or better for an audio file, or the same video as an MP4 with the cleaned track laid back in, so a talking-head clip or a recorded webinar never needs its sound exported separately. Neither carries a watermark, and neither arrives as a public link.

Worth knowing before you start: trimming the dead air is the only step that changes how long the recording runs. If you have already written chapter markers, timestamps or show notes against the original, ask for the pauses to be left alone and the length comes back exactly as it went in.

What to say when you cannot name the problem

You do not need the vocabulary. Describe the sound the way you would describe it to a friend, because the description is the useful part. "There is a buzz behind everything." "My guest is half my volume." "It sounds like we recorded in a stairwell."

If you would rather not decide by ear, upload it and ask what is wrong with the audio. Send the audio as it is, or send the video if that is where the bad sound lives. What comes back is about your specific recording rather than about recordings in general, which is usually the quickest route to knowing which pile yours belongs in.

Saying which problem bothers you still changes the result, because it changes what gets dealt with first. "Take out the air conditioning but keep the voices natural" and "match our levels, the noise does not bother me" are two different jobs on the same file. And the first result is never the last word: "too aggressive, the voice went thin" or "tighten the pauses further" both land on the same recording in the same conversation, with nothing uploaded twice.

What to change before the next recording

None of this helps the file you already have. All four are about the next one.

Leave the loud words some room. Aim the loudest moments about two thirds of the way up the meter instead of at the top. A recording made too quietly can be raised afterwards. A recording that hit the ceiling cannot be unflattened, and it is the one problem with no repair at all.

Get closer. Every step backwards hands more of the take to the room, and that proportion is decided the instant you press record. Nothing downstream moves it, which is why an ordinary microphone near a mouth will beat an expensive one across a table.

Wear one earbud while you record. Almost every unfixable problem announces itself in the first ten seconds if you are listening through the microphone rather than to the room. Your ears filter out the fridge. The headphones do not.

On remote calls, record your own side locally as well. The platform's copy is the one that stutters when the connection dips. A local track keeps your half intact, and half a conversation you can hear is worth more than a whole one you cannot.

Almost every recording sorts itself in one careful listen. Whatever sits under the voice comes off it. A single bang is a cut, not a filter. A room stays a room, because there is no layer to lift off it. Two of those take a sentence to fix, the third takes a different plan, and the only part that has to happen tonight is working out which one you are holding. Send it over and ask if you would rather not be the one deciding.

Related tools

Your next video is one conversation away.

Free account with credits included. No credit card, no learning curve.