Can You Remove Background Music from a Recording?
Yes, if the unwanted sound is music. AI source separation splits a recording into a voice track and everything else, so the music ends up in a separate file. Use the AI Vocal Remover tool (not the MP3 converter on this page) and choose Vocals Only.
- Music, beats, jingles: removed well, because the model was built to separate music.
- Other voices (TV, people nearby): stay in the voice track, because every human voice counts as vocals.
- Steady noise (hum, wind, traffic): only partly removed. Use a noise-reduction tool for that.
How to Remove Background Music
Removing background music from a recording takes three steps. The AI does all the heavy lifting — you just upload, choose the right mode, and download.
- Upload your file to the AI Vocal Remover. Open the AI Vocal Remover tool and drop your recording into it, or click to browse. It accepts MP3, WAV, FLAC, OGG, M4A, AAC, WMA, and MP4 or WebM video, up to 50 MB. Use the highest-quality source file you have — a lossless WAV or FLAC will produce cleaner separation than a compressed MP3.
- Select "Vocals Only" mode. This is the critical step. In Vocals Only mode the Demucs AI splits your audio into two tracks: vocals, which holds all human speech and singing, and an instrumental track with everything else. The background music ends up in the instrumental track, leaving you with clean dialogue.
- Download the vocals track. Once processing completes, download the Vocals file (320 kbps MP3). It contains your speech with the background music removed. You can use it directly or import it into your audio or video editor to replace the original mixed track.
Key point: "Vocals Only" mode keeps all human voices — both the primary speaker and any background voices. If someone is talking on a TV in the background, that speech may remain in the output alongside your primary voice. The AI treats all human vocalization the same way.
When You Need to Remove Background Music
This tool solves a specific problem: you have a recording where the speech is good, but unwanted music is playing in the background. Here are the most common scenarios.
- Podcast cleanup. A guest recorded their side of the conversation with music playing in their room, or a co-host had a Spotify playlist running that bled into their microphone. The speech is perfectly usable, but the background music makes the episode sound unprofessional and creates potential copyright issues. Running the audio through Vocals Only mode strips the music while preserving the conversation.
- Interview recordings. Interviews conducted in cafes, restaurants, or events often pick up background music from the venue's sound system. The interviewee's answers are clear enough to understand, but the ambient music is distracting and makes the recording hard to use in a documentary, news piece, or article. AI separation isolates the voices from the venue soundtrack.
- Video narration with soundtrack. You recorded a voiceover or narration over a video that already had background music baked into the audio track. Now you need the narration without the music — perhaps to re-edit the video with different music, or to use the narration in a different context. Demucs separates the spoken narration from the underlying soundtrack.
- Voiceover extraction from video. A training video, explainer, or presentation has a narrator speaking over background music. You want to reuse the narration in a new project, translate it, or transcribe it accurately. Extracting clean speech without the music makes transcription far more accurate and gives you a usable isolated voiceover track.
- Cleaning up recordings with background TV or radio. Someone recorded a voice memo, phone call, or home video while a TV show, radio station, or music stream was playing in the background. The background audio is distracting and may contain copyrighted content. The AI can remove the musical components, significantly cleaning up the recording.
Speech vs Music Separation
Understanding how the AI separates audio helps you set realistic expectations for the output quality.
Demucs is a deep neural network trained on music. It learned to decompose mixed audio into four stems: vocals (any human voice — singing or speaking), drums (percussion), bass (bass guitar, synth bass, low-frequency instruments), and other (everything else — guitars, keyboards, strings, synths, sound effects). When you select Vocals Only, you get the vocal stem plus one instrumental file that combines the other three.
This means the AI removes all non-vocal sounds, not just "music" in the traditional sense. Here is what gets separated:
- Removed: background music, instrumental loops, soundtrack, jingles, guitar, piano, synthesizers, drum beats, bass lines, ambient music beds.
- Kept: speech, singing, humming, laughter, vocal breaths, lip sounds — anything produced by the human voice.
- Partially removed: ambient noise, room reverb, wind, traffic, air conditioning hum. These non-musical, non-vocal sounds do not fit neatly into any of the four stem categories. The AI handles them inconsistently — some ambient noise ends up in the vocals stem, some in the other stem. You will get a cleaner recording, but do not expect total ambient noise elimination.
The practical takeaway: if your recording has speech mixed with music, the separation will be very effective. If the unwanted sound is non-musical ambient noise (traffic, wind, HVAC), the results will be partial. For pure noise reduction without music separation, a dedicated noise-reduction tool is more appropriate.
Tips for Clean Speech Extraction
The AI does most of the work, but the quality of your input directly affects the quality of the output. Follow these guidelines for the cleanest possible speech extraction.
- Use the highest quality source file. WAV and FLAC files preserve all audio detail, giving the neural network the most information to work with. If you only have an MP3, use the highest bitrate version available. A 320 kbps MP3 will separate better than a 128 kbps version of the same recording because it retains more spectral information that the AI uses to distinguish speech from music.
- Ensure the speech is louder than the music. AI separation works best when the target signal (speech) is the dominant component. Recordings where speech and music are at similar volume levels produce good results. Recordings where music is significantly louder than the speech are harder — the AI may lose some speech detail along with the music. If possible, adjust the mix before processing so the speech sits on top of the music.
- Minimize other noise sources. Background music is what you want to remove, but other noise layers (room echo, wind, hiss) add complexity. The AI handles one separation task very well — splitting vocals from instruments. Adding noise on top of music on top of speech makes all three harder to untangle. Record in a quiet environment when possible, even if music is unavoidable.
- Trim to the relevant section. If only part of your recording has the background music problem, trim the file to that section before uploading. Shorter files process faster and you avoid re-processing sections that are already clean. You can rejoin the segments afterward in any audio editor.
- Check both the vocals and instrumental outputs. Sometimes a small amount of speech leaks into the instrumental stem, or a small amount of music leaks into the vocals stem. Listening to both outputs helps you identify any separation artifacts. If the vocals stem has music bleed, process the file again with Best quality, which uses a fine-tuned model.
Alternative: Extract Audio from Video First
MP4 and WebM videos can go straight into the vocal remover: it reads the audio track and ignores the picture. MOV, AVI and MKV files need an extra step first. Here is the workflow:
- Extract the audio track from your video. Use a tool like FFmpeg (
ffmpeg -i video.mp4 -vn -acodec pcm_s16le audio.wav) or any online video-to-audio converter. Extract as WAV for the best quality. If the video has multiple audio tracks (e.g., narration on track 1, music on track 2), you may already have a clean separation and do not need AI at all — check your video editor's audio track settings first. - Upload the extracted audio to the vocal remover. Select Vocals Only mode and process. The AI will separate the speech from the background music in the extracted audio track.
- Replace the audio in your video editor. Import the cleaned vocal track back into your video editing software (Premiere Pro, DaVinci Resolve, Final Cut Pro, CapCut, or any editor). Mute or delete the original audio track and sync the clean vocals track in its place. Most editors let you snap the new audio to the timeline start position for perfect alignment.
This three-step workflow is standard for video producers who need to clean up interview footage, remove copyrighted music from user-generated content, or isolate narration for re-editing. Extracting the audio first also keeps the upload small, which helps with the 50 MB limit. If you would rather drop the whole soundtrack, our tool to remove audio from video strips it in one step.
Other Ways to Remove Background Music
Forum threads on this question point to the same few options. Plain noise reduction does not work on music, because music changes all the time and noise filters expect a steady sound.
| Tool | Method | Cost | Notes |
|---|---|---|---|
| CleverUtils AI Vocal Remover | AI separation (Demucs) | Free | Audio or MP4/WebM up to 50 MB, runs in the browser |
| Ultimate Vocal Remover | AI separation models | Free, desktop | Runs on your own computer |
| iZotope RX | Music Rebalance / Dialogue Isolate | Paid, desktop | Post-production suite |
| DaVinci Resolve Studio | Voice Isolation | Paid, desktop | Built into the video editor |
| Audacity Noise Reduction | Noise profile | Free, desktop | Does not remove music |