Skip to main content

Remove Background Music from Audio or Video with AI

Remove background music from a voice recording, podcast or video. AI separates speech from the music and gives you a clean voice track as a 320 kbps MP3.

Tap to choose your file

or

Supports M4A, WAV, FLAC, OGG, AAC, WMA, AIFF, OPUS • Max 100 MB

Can You Remove Background Music from a Recording?

Yes, if the unwanted sound is music. AI source separation splits a recording into a voice track and everything else, so the music ends up in a separate file. Use the AI Vocal Remover tool (not the MP3 converter on this page) and choose Vocals Only.

  • Music, beats, jingles: removed well, because the model was built to separate music.
  • Other voices (TV, people nearby): stay in the voice track, because every human voice counts as vocals.
  • Steady noise (hum, wind, traffic): only partly removed. Use a noise-reduction tool for that.

How to Remove Background Music

Removing background music from a recording takes three steps. The AI does all the heavy lifting — you just upload, choose the right mode, and download.

  1. Upload your file to the AI Vocal Remover. Open the AI Vocal Remover tool and drop your recording into it, or click to browse. It accepts MP3, WAV, FLAC, OGG, M4A, AAC, WMA, and MP4 or WebM video, up to 50 MB. Use the highest-quality source file you have — a lossless WAV or FLAC will produce cleaner separation than a compressed MP3.
  2. Select "Vocals Only" mode. This is the critical step. In Vocals Only mode the Demucs AI splits your audio into two tracks: vocals, which holds all human speech and singing, and an instrumental track with everything else. The background music ends up in the instrumental track, leaving you with clean dialogue.
  3. Download the vocals track. Once processing completes, download the Vocals file (320 kbps MP3). It contains your speech with the background music removed. You can use it directly or import it into your audio or video editor to replace the original mixed track.

Key point: "Vocals Only" mode keeps all human voices — both the primary speaker and any background voices. If someone is talking on a TV in the background, that speech may remain in the output alongside your primary voice. The AI treats all human vocalization the same way.

When You Need to Remove Background Music

This tool solves a specific problem: you have a recording where the speech is good, but unwanted music is playing in the background. Here are the most common scenarios.

  • Podcast cleanup. A guest recorded their side of the conversation with music playing in their room, or a co-host had a Spotify playlist running that bled into their microphone. The speech is perfectly usable, but the background music makes the episode sound unprofessional and creates potential copyright issues. Running the audio through Vocals Only mode strips the music while preserving the conversation.
  • Interview recordings. Interviews conducted in cafes, restaurants, or events often pick up background music from the venue's sound system. The interviewee's answers are clear enough to understand, but the ambient music is distracting and makes the recording hard to use in a documentary, news piece, or article. AI separation isolates the voices from the venue soundtrack.
  • Video narration with soundtrack. You recorded a voiceover or narration over a video that already had background music baked into the audio track. Now you need the narration without the music — perhaps to re-edit the video with different music, or to use the narration in a different context. Demucs separates the spoken narration from the underlying soundtrack.
  • Voiceover extraction from video. A training video, explainer, or presentation has a narrator speaking over background music. You want to reuse the narration in a new project, translate it, or transcribe it accurately. Extracting clean speech without the music makes transcription far more accurate and gives you a usable isolated voiceover track.
  • Cleaning up recordings with background TV or radio. Someone recorded a voice memo, phone call, or home video while a TV show, radio station, or music stream was playing in the background. The background audio is distracting and may contain copyrighted content. The AI can remove the musical components, significantly cleaning up the recording.

Speech vs Music Separation

Understanding how the AI separates audio helps you set realistic expectations for the output quality.

Demucs is a deep neural network trained on music. It learned to decompose mixed audio into four stems: vocals (any human voice — singing or speaking), drums (percussion), bass (bass guitar, synth bass, low-frequency instruments), and other (everything else — guitars, keyboards, strings, synths, sound effects). When you select Vocals Only, you get the vocal stem plus one instrumental file that combines the other three.

This means the AI removes all non-vocal sounds, not just "music" in the traditional sense. Here is what gets separated:

  • Removed: background music, instrumental loops, soundtrack, jingles, guitar, piano, synthesizers, drum beats, bass lines, ambient music beds.
  • Kept: speech, singing, humming, laughter, vocal breaths, lip sounds — anything produced by the human voice.
  • Partially removed: ambient noise, room reverb, wind, traffic, air conditioning hum. These non-musical, non-vocal sounds do not fit neatly into any of the four stem categories. The AI handles them inconsistently — some ambient noise ends up in the vocals stem, some in the other stem. You will get a cleaner recording, but do not expect total ambient noise elimination.

The practical takeaway: if your recording has speech mixed with music, the separation will be very effective. If the unwanted sound is non-musical ambient noise (traffic, wind, HVAC), the results will be partial. For pure noise reduction without music separation, a dedicated noise-reduction tool is more appropriate.

Tips for Clean Speech Extraction

The AI does most of the work, but the quality of your input directly affects the quality of the output. Follow these guidelines for the cleanest possible speech extraction.

  • Use the highest quality source file. WAV and FLAC files preserve all audio detail, giving the neural network the most information to work with. If you only have an MP3, use the highest bitrate version available. A 320 kbps MP3 will separate better than a 128 kbps version of the same recording because it retains more spectral information that the AI uses to distinguish speech from music.
  • Ensure the speech is louder than the music. AI separation works best when the target signal (speech) is the dominant component. Recordings where speech and music are at similar volume levels produce good results. Recordings where music is significantly louder than the speech are harder — the AI may lose some speech detail along with the music. If possible, adjust the mix before processing so the speech sits on top of the music.
  • Minimize other noise sources. Background music is what you want to remove, but other noise layers (room echo, wind, hiss) add complexity. The AI handles one separation task very well — splitting vocals from instruments. Adding noise on top of music on top of speech makes all three harder to untangle. Record in a quiet environment when possible, even if music is unavoidable.
  • Trim to the relevant section. If only part of your recording has the background music problem, trim the file to that section before uploading. Shorter files process faster and you avoid re-processing sections that are already clean. You can rejoin the segments afterward in any audio editor.
  • Check both the vocals and instrumental outputs. Sometimes a small amount of speech leaks into the instrumental stem, or a small amount of music leaks into the vocals stem. Listening to both outputs helps you identify any separation artifacts. If the vocals stem has music bleed, process the file again with Best quality, which uses a fine-tuned model.

Alternative: Extract Audio from Video First

MP4 and WebM videos can go straight into the vocal remover: it reads the audio track and ignores the picture. MOV, AVI and MKV files need an extra step first. Here is the workflow:

  1. Extract the audio track from your video. Use a tool like FFmpeg (ffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav) or any online video-to-audio converter. Extract as WAV for the best quality. If the video has multiple audio tracks (e.g., narration on track 1, music on track 2), you may already have a clean separation and do not need AI at all — check your video editor's audio track settings first.
  2. Upload the extracted audio to the vocal remover. Select Vocals Only mode and process. The AI will separate the speech from the background music in the extracted audio track.
  3. Replace the audio in your video editor. Import the cleaned vocal track back into your video editing software (Premiere Pro, DaVinci Resolve, Final Cut Pro, CapCut, or any editor). Mute or delete the original audio track and sync the clean vocals track in its place. Most editors let you snap the new audio to the timeline start position for perfect alignment.

This three-step workflow is standard for video producers who need to clean up interview footage, remove copyrighted music from user-generated content, or isolate narration for re-editing. Extracting the audio first also keeps the upload small, which helps with the 50 MB limit. If you would rather drop the whole soundtrack, our tool to remove audio from video strips it in one step.

Other Ways to Remove Background Music

Forum threads on this question point to the same few options. Plain noise reduction does not work on music, because music changes all the time and noise filters expect a steady sound.

Tool Method Cost Notes
CleverUtils AI Vocal Remover AI separation (Demucs) Free Audio or MP4/WebM up to 50 MB, runs in the browser
Ultimate Vocal Remover AI separation models Free, desktop Runs on your own computer
iZotope RX Music Rebalance / Dialogue Isolate Paid, desktop Post-production suite
DaVinci Resolve Studio Voice Isolation Paid, desktop Built into the video editor
Audacity Noise Reduction Noise profile Free, desktop Does not remove music

Tap to choose your file

or

Supports M4A, WAV, FLAC, OGG, AAC, WMA, AIFF, OPUS • Max 100 MB

Frequently Asked Questions

In most cases, yes. The Demucs AI model separates audio into stems (vocals, drums, bass, other instruments), and the vocals stem contains speech and singing with the music removed. When the music and speech occupy different frequency ranges and do not overlap heavily, the separation is very clean. When speech and music overlap significantly — for example, someone talking over a loud guitar solo in the same frequency range — some musical artifacts may remain, but the speech will still be much clearer than the original.
Partially. Demucs is trained to separate musical stems — vocals, drums, bass, and other instruments. Background TV or radio audio that contains music will be removed effectively. Spoken dialogue from a TV in the background may end up in the vocals stem along with your primary speech, since the model treats all human voices as vocals. For best results, the primary speaker should be louder than any background voices.
Lossless formats like WAV and FLAC give the AI the most data to work with and produce the cleanest separation. MP3 and AAC files work fine but have already lost some audio information during compression, which can slightly reduce separation quality. Avoid heavily compressed files (MP3 at 64 kbps or lower) if possible — the compression artifacts can confuse the separation model. The tool accepts MP3, WAV, FLAC, OGG, M4A, AAC, WMA, and MP4 or WebM video. AIFF is not accepted, so convert it to WAV first.
MP4 and WebM files, yes: upload the video and the tool processes its audio track. The result is an audio file, not a new video, so replace the original audio in your video editor with the cleaned voice track. For MOV, AVI or MKV, extract the audio first, then upload it.
Processing time depends on the length of the file and the quality mode. The tool estimates 1–3 minutes for a typical song-length file in Fast mode and 5–10 minutes in Best mode. Longer files take proportionally longer. For long podcast episodes, split the file into parts: uploads are limited to 50 MB and very long jobs can hit the processing time limit.
The separated speech will sound slightly different from the original because the AI is reconstructing the vocal stem from a mixed signal. In most cases the difference is minimal — the speech is clear, natural-sounding, and free of background music. Occasionally you may notice very subtle artifacts like slight reverb changes or minor tonal shifts in quiet passages. These are generally imperceptible to listeners and far less distracting than the background music that was removed.
Any tool built on AI source separation can do it: this free online vocal remover, the free desktop app Ultimate Vocal Remover, or paid editors such as iZotope RX and DaVinci Resolve Studio. Noise reduction tools are not built for music and usually leave it in.
Not reliably. Its Noise Reduction effect expects a steady noise, and Vocal Reduction and Isolation only works on stereo files with the voice in the center. For music behind speech, AI separation gives cleaner results.

More AI Vocal Remover Guides

Karaoke Maker Online Free — Create Karaoke from Any Song
Make karaoke tracks from any song with AI. Remove vocals and keep the instrumental backing track instantly.
Isolate Vocals from Song Online Free — AI Vocal Extractor
Extract clean vocals from any song with AI. Get isolated vocal tracks for remixes, samples, and covers.
Isolate Drums from Song Online Free — AI Drum Track Extractor
Extract the drum track from any song with AI. Isolate percussion for practice, remixing, or music production.
Acapella Extractor Online Free — Get Vocals from Any Song
Extract acapella from any song with AI. Get clean vocal-only tracks for DJ sets, mashups, and music production.
Back to AI Vocal Remover

Request a Feature

0 / 2000