Skip to main content

Sample Rate Conversion: How the SoXr Resampler Works

Sample rate conversion changes audio from one rate to another, such as 48 kHz to 44.1 kHz. How it works and how our WAV converter uses SoXr.

Convert to WAV with SoXr

SoXr resampling is applied automatically to every conversion

Audio WAV

Tap to choose your file

or

Supports M4A, WAV, FLAC, OGG, AAC, WMA, AIFF, OPUS • Max 100 MB

SoXr resampling applied automatically. Files auto-deleted within 2 hours.

How to Convert 48 kHz to 44.1 kHz

On this page

Upload the file in the converter above and download the WAV. The converter on this page always writes 44.1 kHz, 16-bit stereo, so a 48 kHz source is resampled with SoXr and reduced to 16-bit with Shibata dither.

Audacity

Open the file, choose File → Export Audio, pick WAV and set the sample rate to 44100 Hz in the export dialog.

FFmpeg

The same resampler as this converter: ffmpeg -i input.wav -af aresample=resampler=soxr:precision=28 -ar 44100 output.wav

What is Audio Resampling?

When you convert audio from one sample rate to another (e.g. a 44.1 kHz MP3 → 48 kHz WAV for video editing), every single sample must be recalculated on a new time grid. This process is called resampling.

A naïve approach — simply dropping or duplicating samples — creates audible clicks and aliasing. Professional resamplers use mathematical interpolation (typically polyphase FIR filters) to reconstruct a continuous signal from discrete samples, then re-sample it at the new rate. The quality of this interpolation determines whether your audio stays transparent or picks up artifacts.

Key concept: According to the Nyquist-Shannon theorem, any band-limited signal sampled above twice its highest frequency can be perfectly reconstructed. Resampling leverages this theorem — a high-quality resampler can change rates with zero audible degradation.

What is SoXr?

The SoXr (SoX Resampler Library) is an open-source, audiophile-grade resampling engine originally developed for the SoX (Sound eXchange) command-line audio tool. It uses an FFT-based polyphase algorithm that produces results virtually indistinguishable from the original signal.

SoXr is also available in tools such as SoX, VLC and mpv. CleverUtils.com integrates SoXr through FFmpeg's aresample filter, applying it to every WAV conversion automatically.

Parameter Value What It Does
EngineSoXr (CR64)64-bit double-precision floating-point computation
Precision28-bit~168 dB signal-to-noise ratio — far beyond audible noise floor
DitheringShibataPsychoacoustically-shaped noise that pushes quantization artifacts away from the 1–5 kHz hearing sensitivity peak
Anti-aliasingAutomaticSteep low-pass filter prevents aliasing when downsampling

SoXr vs FFmpeg's Default Resampler

FFmpeg includes two resampling backends: the default swresample (SWR) and the optional soxr. Here's how they compare:

Aspect swresample (default) SoXr
AlgorithmKaiser-windowed sinc (linear phase)FFT-based oversampled polyphase
Internal precisionFollows the sample format in use64-bit double (CR64 engine)
Aliasing rejectionGood at default settingsExcellent (−168 dB with precision=28)
DitheringSeparate FFmpeg option, off unless setShibata (set separately; used by this converter)
SpeedFasterSlower
Passband rippleMeasurable near NyquistNegligible
Best forReal-time streaming, video playbackMastering, archival, distribution

Bottom line: swresample is optimized for speed and is perfectly fine for real-time playback. SoXr is optimized for quality and is the right choice when you're producing a file that will be kept, distributed, or edited further — exactly what a converter does.

Shibata Dithering Explained

When audio is converted between bit depths (e.g. 32-bit float internal processing → 16-bit WAV output), rounding errors create quantization noise. Dithering adds a tiny amount of noise before rounding to eliminate the more unpleasant distortion patterns.

Not all dithering is equal. Standard triangular dithering (TPDF) distributes noise evenly across the frequency spectrum. Shibata dithering uses psychoacoustic noise shaping to push that noise into frequency ranges where human hearing is least sensitive:

Dither Type Noise Distribution Audibility
None (truncation)No noise addedWorst — audible harmonic distortion
Rectangular (RPDF)Flat, randomRemoves distortion, flat noise floor
Triangular (TPDF)Flat, uncorrelatedBetter — no modulation noise
Shibata (noise-shaped)Shifted away from 1–5 kHzLeast audible — exploits hearing curve

Why it matters: Human hearing is most sensitive between 1–5 kHz (the Fletcher-Munson curve). Shibata dithering pushes quantization noise into the less sensitive high-frequency region above 10 kHz, making it effectively inaudible even on high-end monitoring equipment.

When Does Resampling Happen?

SoXr is applied automatically to every WAV conversion on CleverUtils.com, but its impact is most significant in these scenarios:

Scenario Example SoXr Impact
Downsampling hi-res96 kHz FLAC → 44.1 kHz WAVCritical — prevents aliasing artifacts
Music → video rate44.1 kHz MP3 → 48 kHz WAVImportant — clean rate conversion
Voice downsampling48 kHz podcast → 22.05 kHz WAVImportant — preserves speech clarity
Same rate conversion44.1 kHz MP3 → 44.1 kHz WAVMinimal — dithering still applies for bit-depth changes

The biggest quality difference is during downsampling — when the target rate is lower than the source. Without proper anti-aliasing (which SoXr handles automatically), frequencies above the new Nyquist limit fold back into the audible range as distortion.

28-bit Precision: What It Means

SoXr's precision=28 parameter sets the internal computation to 28 effective bits using the CR64 (constant-rate, 64-bit) engine. This translates to approximately 168 dB of signal-to-noise ratio.

For context:

  • 16-bit audio has ~96 dB dynamic range
  • 24-bit audio has ~144 dB dynamic range
  • SoXr at precision=28 computes at ~168 dB — 24 dB below the noise floor of even 24-bit audio

This means the resampling process itself introduces zero audible noise, even for 24-bit masters. The resampler's internal computation is quieter than the quietest sound any real-world recording can capture.

Why not precision=32? Higher precision values increase CPU time with diminishing returns. At precision=28, SoXr already operates 24 dB below 24-bit audio's noise floor — increasing further would make no audible difference.

How CleverUtils Uses SoXr

Every WAV conversion on CleverUtils.com runs through this pipeline:

  1. Upload — your audio file is received over HTTPS
  2. Decode — FFmpeg reads the source format (MP3, FLAC, M4A, OGG, etc.)
  3. Resample — SoXr converts to your chosen sample rate and bit depth
  4. Dither — Shibata noise shaping is applied during bit-depth conversion
  5. Encode — clean PCM samples are written to the WAV container
  6. Download — your WAV file is ready

The entire process is automatic. You just pick your target settings (sample rate, bit depth, channels) and CleverUtils handles the rest using SoXr under the hood. No configuration required, no "quality mode" toggle — every conversion gets the same studio-grade resampling.

Should You Export at 44.1 or 48 kHz?

Destination Sample rate
Music release, CD 44.1 kHz
Video: YouTube, film, broadcast 48 kHz
Podcast 44.1 or 48 kHz; match your editing project
Recording or hi-res archive 96 kHz

Convert once, at the end. Each conversion is a filtering step; a good resampler keeps it inaudible, but going back and forth between rates gains nothing.

Ready to Convert?

Convert your audio to WAV with SoXr resampling

Audio WAV

Tap to choose your file

or

Supports M4A, WAV, FLAC, OGG, AAC, WMA, AIFF, OPUS • Max 100 MB

Frequently Asked Questions

SoXr (SoX Resampler Library) is an audiophile-grade resampling engine that uses FFT-based polyphase algorithms at 28-bit precision. FFmpeg's default swresample uses a Kaiser-windowed sinc filter that is fast and good, with more measurable aliasing near the Nyquist frequency at default settings. SoXr pushes these artifacts far below audibility. The Shibata dithering used here is a separate FFmpeg option.

Shibata dithering is a psychoacoustically-optimized noise shaping method that pushes quantization noise away from the 1–5 kHz range where human hearing is most sensitive (the Fletcher-Munson curve). The result is dither noise that is less perceptible than standard triangular or rectangular dithering, even though the total noise energy is similar.

For most casual listening, the difference is inaudible. SoXr matters most when downsampling hi-res audio (e.g. 96 kHz to 44.1 kHz) where aliasing from a lesser resampler could become audible on high-end monitoring equipment. It also ensures bit-perfect transparency for professional mastering workflows where cumulative processing errors matter.

No. SoXr is applied automatically to every WAV conversion on CleverUtils.com. Simply upload your file, choose your WAV settings (sample rate, bit depth, channels), and the SoXr resampler handles the rest. No special toggle or configuration needed.

Resampling occurs whenever the target sample rate differs from the source — for example, converting a 44.1 kHz MP3 to 48 kHz WAV for video, or downsampling a 96 kHz FLAC to 44.1 kHz for CD burning. Even when sample rates match, SoXr's Shibata dithering ensures clean bit-depth transitions (e.g. internal 32-bit float → 16-bit output).

Neither sounds better for listening: both capture the full range of hearing (up to 22.05 and 24 kHz). Pick 44.1 kHz for music and CD, 48 kHz for video.

Not for playback in practice. 192 kHz captures frequencies up to 96 kHz, far above hearing, and the file is 4x larger than at 48 kHz. Higher rates can help inside some production steps, not in the final file.

96,000 samples per second, which captures frequencies up to 48 kHz. It is used for recording and hi-res releases. A 96 kHz, 24-bit stereo WAV takes about 34.6 MB per minute.

More MP3 to WAV Guides

WAV Sample Rate & Bit Depth Explained: Which Settings to Use
44.1 vs 48 kHz, 16-bit vs 24-bit, stereo vs mono — which WAV settings should you use?
MP3 to WAV Speed Changer: Adjust Tempo for Editing
Slow down MP3 for transcription or speed up for editing. Uncompressed WAV output for DAW compatibility.
MP3 to WAV Bass Boost: Uncompressed Output for Speakers
Bass boost without lossy re-encoding. Perfect for car audio, PA systems, and DJ setups.
MP3 to WAV Volume Boost: Amplify for DAW Editing
Boost quiet MP3 files by +3 to +20 dB and convert to WAV for editing in Audacity, Logic, or Reaper.
MP3 to WAV Fade In/Out: Uncompressed Output with Smooth Transitions
Add fade in and fade out to MP3 files and convert to WAV. Pre-faded output for DAW editing.
Normalize MP3 to WAV Loudness: Consistent Volume for DJs, Editors & Playlists
Normalize MP3 files to consistent WAV volume. Eliminate volume jumps between tracks from different albums and eras.
Does Converting MP3 to WAV Improve Quality? (Myth Busted)
Converting MP3 to WAV does NOT restore lost data. Why the file gets bigger without improving quality, and when conversion still makes sense.
Back to MP3 to WAV Converter

Request a Feature

0 / 2000