Skip to main content

Audio Editing Basics

Cuts, fades, crossfades, level matching, room tone, EQ, compression, and LUFS export standards for UK broadcast journalists and podcasters.

Last reviewed: Next review due:

Audio editing in broadcast journalism

Audio editing is the craft of assembling recorded sounds into a broadcast-quality package or podcast episode. It involves making cuts, adjusting levels, filling gaps, and ensuring the final output is clean, intelligible, and meets technical loudness standards. Even journalists who work primarily in print or online need basic audio editing skills for video content, social media clips, and podcast production.

The most important principle in audio editing is the same as in writing: everything that does not serve the story should be removed. An over-long interview clip, an awkward pause, a distracting background sound — all of these can and should be edited out. The listener should not be aware of the editing; the package should feel natural and continuous.

UK broadcast standard tools include Hindenburg Journalist Pro (the most widely used radio/podcast editor among working UK journalists), Adobe Audition, and Descript. BBC journalists use ProTools and Hindenburg. Audacity is free and functional as a learning tool.

Core audio editing techniques

Hard cuts

The most basic edit: one piece of audio ends and another begins at an exact point. Works well between sections with consistent background noise. Abrupt cuts in inconsistent acoustic environments sound “clicky.”

Fades

A gradual increase or decrease in volume at the start or end of a clip. Use short fades (0.05–0.1 sec) at edit points to avoid clicks. Longer fades (0.5–2 sec) for music beds or transitions.

Crossfades

Overlapping fade-out and fade-in between two clips. Used to smooth transitions between sections with different background noise levels. Hindenburg and Audition both handle crossfades automatically.

Level matching

Ensuring all clips in a package are at a consistent volume. Check each clip individually: aim for speech averaging around –18 to –12 dB RMS, peaking no higher than –6 dBFS.

Room tone fill

Using recorded room tone (ambient silence from the recording location) to fill gaps between edit cuts. Without it, edit points sound hollow or “dead.”

EQ

High-pass filter at 80–100 Hz removes room rumble. Gentle boost at 2–4 kHz adds vocal presence. Cut any resonant room frequencies that sound boxy.

Compression

Reduces dynamic range: loud parts quieter, quiet parts louder. For speech: 3:1 ratio, fast attack (5–10 ms), medium release (50–100 ms), 6–8 dB of gain reduction maximum.

LUFS normalisation

Integrated loudness target for podcast: –16 LUFS. True peak ceiling: –1 dBTP. Use a loudness meter (Youlean, LUFS Meter, or built-in in Hindenburg/Audition) before export.

Red flags in audio editing

  • Editing in MP3 format — always work in WAV; only export to MP3 as the final step.
  • No room tone recorded — edit points between clips with inconsistent backgrounds will sound obvious.
  • Clipping (audio peaks above 0 dBFS) — causes digital distortion that cannot be removed in post.
  • Over-compressing speech — heavy compression produces a pumping effect and sounds unnatural.
  • Not saving the raw, unedited session before editing — always keep the original recording.
  • Removing all breaths from speech — makes it sound robotic; reduce, do not eliminate.
  • Exporting above –14 LUFS integrated — platforms will turn it down; there is no benefit.

Audio editing checklist

  • Raw, unedited recording is saved as a backup before any editing begins.
  • All clips are selected and logged with in-words, out-words, and duration noted.
  • Working file is in WAV format (44.1 kHz, 24-bit) — not MP3.
  • All edit points have short fades to avoid clicks (0.05–0.1 sec minimum).
  • Level matching: all speech clips are at a consistent level (–18 to –12 dB RMS average).
  • Room tone is used to fill all gaps between edit cuts.
  • High-pass filter applied at 80–100 Hz to remove low-frequency rumble from all speech tracks.
  • Compression applied lightly to even out dynamic variation in speech.
  • Loudness meter confirms final mix is at –16 LUFS integrated and below –1 dBTP true peak.
  • Final export is to MP3 at 128 kbps minimum (or WAV for broadcast delivery).

Plan before you record

Good editing starts with good recording. Use our pitch generator to plan your audio report structure before you record.

Open Pitch Generator

Common mistakes

  • Not recording room tone — the most common and most avoidable edit-quality problem.
  • Over-editing clips to remove all pauses — speech loses its natural rhythm.
  • Not checking the final export loudness — masters that are too quiet or too loud sound unprofessional.
  • Using music that is not licensed for commercial use without checking clearance.
  • Not labelling and organising session files — returning to an unlabelled session weeks later wastes hours.
  • Applying heavy noise reduction to audio with significant noise problems — the result often sounds worse than the original; re-record if possible.

Related guides

Primary sources

Frequently asked questions

What LUFS target should I use for podcast export?
The widely accepted podcast standard is –16 LUFS integrated loudness (for stereo) with a true peak ceiling of –1 dBTP. Spotify normalises to –14 LUFS; Apple Podcasts accepts up to –16 LUFS. Exporting at –16 LUFS means platforms will leave your audio at its natural level or turn it up slightly — never clipped. Do not master above –14 LUFS; louder masters will be turned down by platforms, wasting any perceived volume advantage.
What is room tone and how do I use it in editing?
Room tone is the ambient sound of your recording location, recorded silently for 30–60 seconds after your interview. In editing, you use room tone to fill gaps between edit cuts — replacing silence or abrupt jumps in background noise with a consistent ambient bed that makes edits inaudible. Without room tone, edit points often produce a clicking or stuttering effect as the background noise level changes suddenly. Always record room tone at every location before you leave.
How do I remove breaths and ums from an audio recording without it sounding unnatural?
Breaths and filler sounds should be removed selectively, not wholesale. Removing every breath makes speech sound robotic; removing obviously intrusive breaths or long ums improves clarity without sounding edited. Technique: use a gentle fade to reduce rather than cut breaths entirely — reduce them by 6–12 dB rather than silencing them completely. For ums and ers: delete the filler and replace the gap with room tone. Descript’s “remove filler words” function does this automatically but should be reviewed manually.
What is the difference between EQ and compression in audio production?
EQ (equalisation) adjusts the balance of frequency content in a recording — boosting or cutting specific frequency ranges. For spoken-word audio, typical EQ moves include: a high-pass filter at 80–100 Hz to remove low-frequency rumble; a gentle boost around 2–4 kHz for vocal presence and clarity; and a cut at any resonant room frequency that sounds boxy. Compression reduces the dynamic range of a recording — making loud parts quieter and quiet parts louder, so the overall level is more consistent. For broadcast speech, light compression (3:1 ratio, fast attack, medium release) produces a natural result.
Should I save sessions in WAV or MP3 during editing?
Always edit in an uncompressed or lossless format: WAV (PCM, 44.1 kHz, 24-bit) is the standard for broadcast audio editing. Never edit in MP3, which is a lossy compressed format — repeated export and re-import of MP3 files degrades audio quality each time (generation loss). Export to MP3 only as the final delivery step: 128 kbps minimum, 192 kbps for music-heavy content, 320 kbps for quality-sensitive productions.