Stereo to Mono: How to Downmix Without Losing Anything
2 October 2026

Stereo to Mono: How to Downmix Without Losing Anything

You export a two-mic interview, the platform asks for a mono track, and one command later the file has gone quiet, hollow, or almost silent. I hit that last one with a podcast recorded in two takes, one of them flipped in polarity, and the merged file was nothing but hiss. The fix isn’t a better converter. It’s knowing what a stereo to mono downmix does to the samples, and checking the source first.

Here’s what this guide covers: the arithmetic behind the downmix, the one failure mode that eats a recording, a two-command test to run first, the stereo to mono recipes worth using and the one that clips half your samples, what changes with a 5.1 source, and how to get the loudness back. Every number below comes from tests on this machine with FFmpeg 7.1.5, so you can repeat them.

What a stereo to mono downmix does to your samples

A stereo file holds two sample streams, a mono file holds one, and the conversion has to invent that single stream. FFmpeg’s -ac 1 flag does it by averaging: half of the left channel plus half of the right, sample by sample. I wrote constant values into a stereo file and read the mono result back sample by sample:

Left sample Right sample Mono result
10000 30000 20000
8000 0 4000
30000 30000 30000
30000 -30000 0
32767 -32768 0

That last row matters more than it looks. Two full-scale samples of opposite polarity average to digital zero. Nothing errors, nothing warns, and the stereo to mono result is silent at exactly the spot where your audio used to be.

The matrix is normalised to sum to one, at 0.5 plus 0.5, because adding the channels would push a loud mix past full scale. FFmpeg’s resampler documentation describes that matrix. That one choice explains every surprise in this article. Identical channels survive at full level. Unrelated channels each arrive at half level, which is 6.02 dB down. Opposed channels cancel. A stereo to mono downmix is exact, lossy, or destructive depending on what your two channels are holding.

The failure mode that eats a recording

Polarity inversions are common in the wild: a mis-wired cable, a plugin with the phase button engaged, a channel pasted back inverted. You won’t hear it in stereo, because your speakers reproduce each side separately. You hear it the moment the stereo to mono downmix runs.

I built a 1 kHz stereo tone at an amplitude of 20000 and ran four variants through -ac 1, measuring RMS from the output file:

Source Mono RMS Mono peak Versus one channel
R identical to L -7.30 dBFS 20000 +0.00 dB
R inverted (R = -L) -330 dBFS 0 -323 dB
R at half level, inverted -19.34 dBFS 5000 -12.04 dB
R silent -13.32 dBFS 10000 -6.02 dB

A perfect inverse pair returns a file of zeros. The peak is 0 and the RMS reads -330 dBFS. If that had been a 40-minute interview, the master would still be fine and the delivered file wouldn’t play.

Test the source before you convert

Phase check before a stereo to mono downmix, comparing L+R and L-R levels with volumedetect
The two-command check: L+R shows what survives, L-R shows what the downmix throws away.

Ten seconds of work tells you what a stereo to mono downmix will do to your file. Sum the channels, then subtract one from the other, and read the levels with volumedetect:

ffmpeg -hide_banner -i input.mp4 -af "pan=mono|c0=c0+c1,volumedetect" -f null -
ffmpeg -hide_banner -i input.mp4 -af "pan=mono|c0=c0-c1,volumedetect" -f null -

The first command measures what the mono file would contain. The second measures what a stereo to mono downmix throws away. I ran that pair on three synthetic sources:

Source L+R mean L-R mean Gap
Dual mono (L equals R) -2.50 dBFS -91.00 dBFS 88.5 dB
One channel inverted -91.00 dBFS -2.50 dBFS 88.5 dB
True stereo (decorrelated) -19.30 dBFS -19.40 dBFS 0.1 dB

Read it like this. If L+R is far louder than L-R, your channels are effectively one signal and the stereo to mono downmix is safe. If L-R wins, some of your audio runs in opposite phase and the downmix will cancel it, so fix the phase or take a single channel. If the two numbers sit within a dB or two, the file is genuinely stereo, and folding it costs about 6 dB. The filter reports at the info log level, so don’t wrap it in -v error or you’ll get an empty log and think the test failed.

Four stereo to mono recipes, and what each one costs

The pan filter gives you the matrix directly. These are measured on a stereo file holding 20000 in the left and 30000 in the right:

Recipe What it does Mono sample
-ac 1 Averages both channels 25000
pan=mono|c0=0.5*c0+0.5*c1 Same average, written out 25000
pan=mono|c0=c0 Keeps the left, discards the right 20000
pan=mono|c0=c0+c1 Raw sum, no attenuation 32767, clipped

Avoid the raw sum. On a full-scale pair it pushed every one of 4800 test samples to the ceiling. On a 700 Hz tone at 24000 amplitude, the average came out at -2.70 dBFS with zero clipped samples, while the unattenuated sum hit 0.00 dBFS and clipped 25000 of 48000 samples, which is 52% of the file. The pan filter’s matrix set is documented in the FFmpeg filter reference.

Keeping one channel is the right call when the sides disagree in phase, or when one of them is a room mic you’d rather lose.

The flag that gets ignored

Add -c:a copy next to -ac 1 and FFmpeg copies the audio stream bit for bit, so the downmix never happens. My test returned exit code 0 and the output still reported two channels. Probe the result instead of trusting the code. If you’d rather work on the audio alone, the audio extraction guide covers that path:

ffprobe -v error -show_entries stream=codec_name,channels,channel_layout -of default=nw=1 output.mp4

What changes with a 5.1 source

A 5.1 source needs a real matrix for a stereo to mono downmix, and FFmpeg builds one. I measured its coefficients by feeding a 1 kHz tone into one channel of a 5.1 file at a time and reading the mono gain:

Channel Measured gain Coefficient
Front left and front right -13.68 dB 0.2071
Centre -10.67 dB 0.2929
LFE silent 0.0000
Back left and back right -16.69 dB 0.1465

Those numbers look arbitrary until you add them up: 0.2071 + 0.2071 + 0.2929 + 0.1465 + 0.1465 = 1.0000. The whole matrix is normalised so the coefficients sum to one, exactly like the stereo case, which keeps a signal present in every channel at unity gain. A tone fed to all six channels came out at +0.00 dB. It also means the LFE, your subwoofer channel, is dropped completely.

The cost appears when the channels carry different content. Six different tones spread across six channels measured -5.9 LUFS as a 5.1 file and -20.5 LUFS after the stereo to mono downmix, a drop of 14.6 LU. The normalisation is doing what it was designed to do. It just assumes your channels are correlated, and a real mix isn’t.

How much loudness a stereo to mono downmix takes

Loudness lost to a stereo to mono downmix, measured at 3 LU for dual mono and 6 LU for true stereo
EBU R128 loudness before and after the downmix, on sources with identical per-channel level.

The loss depends on how similar the channels are, so a single “mono is 3 dB quieter” rule is wrong. I measured integrated loudness with EBU R128 on two stereo files with identical per-channel level, one dual mono and one decorrelated:

Source As stereo After -ac 1
Dual mono (L equals R) -15.9 LUFS -18.9 LUFS
Decorrelated stereo -15.9 LUFS -21.9 LUFS
One channel alone -18.9 LUFS –

Both stereo files measure -15.9 LUFS, because R128 sums channel power and doesn’t care about correlation. The stereo to mono step costs 3 LU on the dual mono file and 6 LU on the true stereo one. The third row is the sanity check: the dual mono downmix lands on the loudness of a single channel.

Don’t guess the compensation. A flat volume=6dB restored my decorrelated test to -15.9 LUFS, but on a dual mono source the same 6 dB would overshoot by 3 dB and start clipping. Use a loudness normaliser with a true-peak ceiling, the same approach as the audio normalisation workflow:

ffmpeg -i input.mp4 -c:v copy -ac 1 -af "loudnorm=I=-16:TP=-1.5:LRA=11" -ar 48000 -c:a aac -b:a 96k output.mp4

That recipe measured -16.2 LUFS at a -5.3 dBFS peak on my decorrelated source.

Pin the sample rate after a stereo to mono downmix

One trap hides in that command. Run loudnorm without the -ar flag and FFmpeg hands back audio at 192000 Hz, because the filter works at a high internal rate and the output keeps it. Six seconds of mono PCM measured 2310246 bytes at 192 kHz against 577614 bytes at 48 kHz, a 4x difference for identical audio. Encoded to AAC, the stream still reported 96000 Hz rather than the 48000 Hz I fed in. The sample rate guide goes deeper on why the number matters.

Does a stereo to mono downmix touch the picture?

No, provided you copy the video stream instead of re-encoding it. I hashed the video stream before and after a stereo to mono downmix and both came back MD5=e1523ba3a655c4da510e350f3d633875. The audio went from two channels to one and the H.264 stream stayed byte-identical, which is what -c:v copy is for. The same stream-copy logic drives the remux method. Skip that flag and FFmpeg re-encodes the picture, so you’ll lose a generation of quality for a change that only concerned the audio.

Delivery size doesn’t drop the way people expect. Encoded to AAC at a 128k target, my 30-second test file was 486782 bytes as stereo and 487025 bytes as mono, a difference of 243 bytes. Lower the target to 64k for the mono file and you get 247530 bytes, about half. If your bottleneck is a platform ceiling rather than bitrate, the upload limit guide covers that fight.

Common mistakes

  • Re-encoding the video while downmixing the audio. Add -c:v copy so the picture survives.
  • Using pan=mono|c0=c0+c1 without attenuation, which clipped 52% of the samples in my test.
  • Trusting exit code 0. With -c:a copy in the same command, -ac 1 is ignored and the file keeps two channels.
  • Compensating with a fixed volume=6dB, which is right for decorrelated stereo and 3 dB too loud for dual mono.
  • Forgetting -ar after loudnorm, which leaves a 192 kHz file four times bigger than it needs to be.
  • Converting to stereo to mono before checking for an inverted channel, the one case where the audio does disappear.

FAQ

Will a stereo to mono downmix make my file smaller?

Not at the same bitrate target. My stereo and mono AAC files at 128k differed by 243 bytes. Drop the target to 64k for the mono version and the file halves to 247530 bytes. Mono buys you quality at a lower bitrate, not a smaller file at the same one.

Can I undo a stereo to mono downmix later?

No. Averaging two channels into one throws the difference away, and you can’t rebuild it later. Keep the stereo master and deliver the mono file as a copy.

Will a mono file play out of both speakers?

Yes, players duplicate the single channel after a stereo to mono downmix. To write that duplication into the file, use pan=stereo|c0=c0|c1=c0, which measured at unity gain: a 20000 sample came back as 20000 in both channels. -ac 2 also duplicates, but it attenuated the same sample to 14142, which is 3.01 dB down.

Does a stereo to mono downmix change the video stream?

Not with -c:v copy. The video MD5 stayed identical in my test while the audio went from stereo to mono. Without that flag, FFmpeg re-encodes the picture as well.

Conclusion

A stereo to mono downmix is one line of arithmetic, and every surprise comes from the source rather than the tool. Identical channels survive intact, unrelated channels lose 6 dB, and opposed channels disappear. So run the L+R and L-R check first, choose between the average, a single channel, or the explicit pan form, put loudnorm after the downmix with the sample rate pinned, and keep -c:v copy in the command so the picture is never touched twice.

It took one ruined podcast to teach me the check was worth ten seconds. Now it’s the first thing I run.