What a head does to sound

Sound arrives at the ear from a direction, and by the time it reaches the eardrum it has been filtered by the shape of the head, the pinnae and the torso. A sound from the left reaches the left ear directly and the right ear indirectly, having travelled further and been shadowed and reflected on the way. Those differences in level, timing and spectrum are the entire basis of spatial hearing.

A head-related transfer function is a measurement of exactly that: the filtering applied to a sound arriving from one direction, captured as the difference between what arrives at each ear. Do that measurement across a grid of directions and you have a dataset capable of placing a sound anywhere in a listener's head.

The realism of spatial audio comes from reproducing a real filtering chain accurately, not from adding channels.

Two different techniques, one name

Binaural recording is the direct route. A microphone pair sits at the entrance to each ear canal of a real person or a manufactured dummy head, and the head does the spatialising. It is simple, it is immediate, and whatever is around the microphones during the recording is captured too, which is exactly what makes it convincing.

Binaural rendering is the indirect route. You record normally in whatever format suits the project, then at playback convolve each channel through the head-related filters for the direction you want that source to appear to come from. It is more flexible because you can change the virtual source position after the fact, and it is also only ever an approximation of a real measurement.

Abstract head-shaped imagery representing head-related transfer functions and their role in spatial playback
Binaural audio works by reproducing the filtering a real head applies to sound arriving from a given direction. The realism depends on those filters being measured correctly.
  • Binaural recording: spatialisation happens in the room, at the microphone, and cannot be undone.
  • Binaural rendering: spatialisation happens at playback, and the mix stays conventional until the end.
  • Both produce two channels intended for headphones, for the same underlying reason.

The trade-offs that do not go away

There are three structural costs, and none of them are fixable with a better plugin. The first is channel count. Binaural is inherently two channels, so it cannot carry a wide sound field the way an array format can, and any decision about width has to be made permanently.

The second is mono. Sum a binaural pair to mono and the spatial information cancels, because that information exists only in the difference between the channels. What survives is a filtered version of the original source. That is not a degraded fold-down, it is a different and much smaller thing.

FormatChannelsNeeds headphonesSurvives mono
True binaural recording2YesNo
Binaural rendering2YesNo
Stereo2NoYes
Ambisonics4+NoPartly

The third is playback. The technique depends on each ear hearing only what it would hear from a real source, and speakers in a room cannot deliver that without crosstalk. On a pair of speakers the left channel reaches both ears and the illusion collapses. Headphones are not a preference here, they are a requirement.

When it is worth it

Binaural is a strong choice for headphone-first delivery, for virtual and augmented reality where the listener supplies their own headphones, and for ambience and Foley where placing a sound behind a listener matters more than reproducing an accurate front stage. It is a poor choice for a mix that has to survive a club system, a phone speaker, or a mono fold-down.

The decision worth making early is the delivery target. Choosing it after the mix is finished means either compromising the mix or rendering everything twice. Choosing it first means the mix is built for the format it will actually be heard in, which is the far cheaper problem.

The bottom line

Binaural works because a real head and torso filter sound arriving from different directions, and reproducing those filters accurately makes a two-channel recording sound placed. The trade-offs are structural rather than fixable: it is two channels, it needs headphones, and it collapses in mono. Use it when the delivery is specifically headphone listening, and use a different format when it is not.

Check the low end survives the fold

Retuning a whole production shifts every channel together, so the level relationship between a centre element and its surrounds survives the move unchanged.

Explore 432Hz MASTER

Frequently asked questions

What is the difference between binaural recording and binaural rendering?

Recording means putting microphones where a person's ears would be, on a real head or a dummy one, and capturing the filtering directly. Rendering means keeping a normal multichannel recording and convolving it through measured head-related filters at playback time. The result can sound similar and the workflow is completely different.

Does binaural audio work on speakers?

Not properly. The technique depends on each ear hearing only what it would hear from a real source direction, which speakers in a room cannot reproduce without crosstalk. Played over speakers the left channel reaches both ears and the illusion breaks down.

Why does binaural collapse in mono?

Because the entire effect lives in the difference between the two channels. Sum them to mono and what is left is a filtered version of the source, with no spatial information at all. A mono fold-down of binaural is not a degraded version, it is a different and much smaller thing.

Is binaural the same as surround?

No. Surround formats place channels in a physical array across a listening area, while binaural aims at a single point inside a head. Ambisonics sits between them, carrying a sound field that can be rendered to either.

More from the atixUniverse blog