• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

Dan Clark E3 Headphone Review

Rate this headphone:

  • 1. Poor (headless panther)

    Votes: 5 1.7%
  • 2. Not terrible (postman panther)

    Votes: 11 3.7%
  • 3. Fine (happy panther)

    Votes: 47 15.9%
  • 4. Great (golfing panther)

    Votes: 233 78.7%

  • Total voters
    296
Why does the E3 sound spatially narrow despite measuring so well?

The E3 has extremely low distortion, an excellent frequency response, clean group delay, and apparently very good time-domain behavior. Yet both my 10-year-old son and I had the same impression with Yosi Horikawa's Bubbles: many of the balls seemed to fall into almost the same spatial location. I have also seen the E3 described as having a “giant IEM” type of presentation.

That seemed difficult to reconcile with the conventional measurements, so I built a simple 3D FDTD simulation, using Claude Fable 5 extensively to help with the implementation and analysis, to look at the sound field across the pinna rather than only at the ear-canal response. The models are deliberately simplified, and the AMTS geometry is approximated from photographs, so this is not intended as an accurate simulation of the commercial headphones.

Figure 1 — Simplified geometries and probe positions
View attachment 552264


Figure 2 — Incident level distribution across the pinna
View attachment 552265

The interesting part is that simple level variation does not seem to explain the difference. The DX10000CL-type model also has substantial spatial level variation.
When the overall level difference at each probe is removed and only the spectral shape is compared, the result becomes much clearer.


Figure 3 — All probe frequency responses. Top: including level differences. Bottom: spectral shape only.

View attachment 552266

Mean spatial variation of spectral shape from 5–10 kHz:
  • MDR-Z1R-type: 0.8 dB
  • LCD-type: 1.5 dB
  • DX10000CL-type: 1.3 dB
  • DCA/AMTS-like: 2.7 dB
I also separated smooth spatial gradients from local irregularity using a local planar fit.

Figure 4 — Local spatial gradient vs. residual irregularity
View attachment 552267


The AMTS-like model is a clear outlier in the 5–10 kHz spectral-shape result.
So my current hypothesis is simply this:

A headphone can have an excellent frequency response at the ear canal while still producing a highly position-dependent spectral field across the pinna.

If so, conventional single-point measurements may be missing a variable relevant to spatial perception. The “giant IEM” impression would also make some intuitive sense: the final FR can be excellent, while the pinna is no longer being illuminated by a spatially consistent wavefront.

This obviously does not prove that AMTS causes the E3's spatial presentation, but I think it is an interesting result worth testing experimentally.
I don't think I understand this, what matters is measured frequency response at the eardrum, which is what for instance Amir's & Oratory's GRAS devices do. What matters is the frequency response that gets down into your ear, not what is measured elsewhere over your pinna. Is that not so, I don't understand the overall mechanism that you're suggesting?
 
I don't think I understand this, what matters is measured frequency response at the eardrum, which is what for instance Amir's & Oratory's GRAS devices do. What matters is the frequency response that gets down into your ear, not what is measured elsewhere over your pinna. Is that not so, I don't understand the overall mechanism that you're suggesting?
Yes, I agree that the frequency response at the eardrum is ultimately what the listener receives, and that is what GRAS-based measurements are designed to capture.

What I am suggesting is slightly different: a single eardrum frequency-response measurement does not necessarily tell us whether the pinna is being acoustically excited in the same way as it would be in a natural sound field.

The pinna creates direction-dependent peaks and notches, particularly in roughly the 4–12 kHz region, and those spectral changes are important for elevation and front/back spatial perception. If the sound field around the pinna is altered—for example by AMTS—the pinna may receive a degraded or less spatially appropriate wavefront.

In that case, it may still be possible to obtain a very good overall frequency response at the eardrum for a particular measurement condition, while the direction-dependent pinna cues needed for spatial perception are not reproduced as accurately.

So I am not arguing that measurements away from the eardrum are more important than the eardrum response. The hypothesis is that a good eardrum FR alone may not fully characterize how well the system preserves the spatially dependent acoustic information that the pinna normally converts into spectral cues.

That is the possible explanation I am considering for why the E3 could measure very well in conventional frequency response, yet still sound spatially worse.
 
Yes, I agree that the frequency response at the eardrum is ultimately what the listener receives, and that is what GRAS-based measurements are designed to capture.

What I am suggesting is slightly different: a single eardrum frequency-response measurement does not necessarily tell us whether the pinna is being acoustically excited in the same way as it would be in a natural sound field.

The pinna creates direction-dependent peaks and notches, particularly in roughly the 4–12 kHz region, and those spectral changes are important for elevation and front/back spatial perception. If the sound field around the pinna is altered—for example by AMTS—the pinna may receive a degraded or less spatially appropriate wavefront.

In that case, it may still be possible to obtain a very good overall frequency response at the eardrum for a particular measurement condition, while the direction-dependent pinna cues needed for spatial perception are not reproduced as accurately.

So I am not arguing that measurements away from the eardrum are more important than the eardrum response. The hypothesis is that a good eardrum FR alone may not fully characterize how well the system preserves the spatially dependent acoustic information that the pinna normally converts into spectral cues.

That is the possible explanation I am considering for why the E3 could measure very well in conventional frequency response, yet still sound spatially worse.
It's a stereo headphone, there is no elevation or positional information beyond stereo and whatever baked in effects might be included in the music. It doesn't matter how the headphone achieves it's frequency response at your eardrum. If two headphones achieved the same frequency response at your eardrum but did it in "different ways" through different physical design choices then it would still sound the same to you.
 
Yes, I agree that the frequency response at the eardrum is ultimately what the listener receives, and that is what GRAS-based measurements are designed to capture.

What I am suggesting is slightly different: a single eardrum frequency-response measurement does not necessarily tell us whether the pinna is being acoustically excited in the same way as it would be in a natural sound field.

The pinna creates direction-dependent peaks and notches, particularly in roughly the 4–12 kHz region, and those spectral changes are important for elevation and front/back spatial perception. If the sound field around the pinna is altered—for example by AMTS—the pinna may receive a degraded or less spatially appropriate wavefront.

In that case, it may still be possible to obtain a very good overall frequency response at the eardrum for a particular measurement condition, while the direction-dependent pinna cues needed for spatial perception are not reproduced as accurately.

So I am not arguing that measurements away from the eardrum are more important than the eardrum response. The hypothesis is that a good eardrum FR alone may not fully characterize how well the system preserves the spatially dependent acoustic information that the pinna normally converts into spectral cues.

That is the possible explanation I am considering for why the E3 could measure very well in conventional frequency response, yet still sound spatially worse.

The E3 is very position-dependent. Moving it just a few millimeters forward or backward on my head can significantly change the response above 4 kHz. For example, play a 5 kHz test tone and move the cups slightly - you can get anything from a deep null to a peak depending on the position and getting both ears positioned exactly the same is extremely difficult.

The E3 can sound great, but the positional variation and seal sensitivity make it less plug-and-play than I would like.
 
It's a stereo headphone, there is no elevation or positional information beyond stereo and whatever baked in effects might be included in the music. It doesn't matter how the headphone achieves it's frequency response at your eardrum. If two headphones achieved the same frequency response at your eardrum but did it in "different ways" through different physical design choices then it would still sound the same to you.
Is that so? Do we have research that backs up this claim? I think we dont know enough about how we perceive headphone sound to say this.
 
Is that so? Do we have research that backs up this claim? I think we dont know enough about how we perceive headphone sound to say this.

What's wrong with the claim? If two headphones truly produced the same FR at both eardrums - not merely the same smoothed magnitude FR - then yes, there would be no acoustic information available to distinguish them. That's basically true by definition.
 
This is interesting but I think it may have more to do with cross feed across the frequency. If you have Roon or a McIntosh MHA100 or 140 with the HxO, turn it off and on and things start to snap to the middle and where instruments sounded hard left or right become more front left or front right.
 
It's a stereo headphone, there is no elevation or positional information beyond stereo and whatever baked in effects might be included in the music. It doesn't matter how the headphone achieves it's frequency response at your eardrum. If two headphones achieved the same frequency response at your eardrum but did it in "different ways" through different physical design choices then it would still sound the same to you.
I agree with the limiting case: if two headphones produced exactly the same pressure waveform at both eardrums, including magnitude and temporal/phase behaviour, then the different physical mechanism used to produce it would be irrelevant.

But that is not quite the hypothesis I am testing.

In my simplified simulation, the AMTS-type geometry produces a much less spatially uniform near-field over the pinna, especially in roughly the 5–10 kHz region. More importantly, the pattern changes substantially with frequency rather than behaving like a simple smooth level gradient.

I am not claiming that the auditory system somehow “reads” pressure at different points on the pinna directly. The question is whether a highly non-uniform, frequency-dependent incident field makes the resulting headphone-to-eardrum transfer function more sensitive to the listener's actual pinna geometry and to small changes in fit or position.

If so, two headphones could measure similarly at the eardrum of one fixed GRAS configuration while behaving less similarly across real human ears or small positional changes.

And that distinction matters because the pinna-related spectral filtering used in spatial perception occurs mainly in this high-frequency region.

So the hypothesis is not “the headphone creates elevation information from stereo.” It is that a spatially non-uniform incident field may interact with the pinna less robustly, potentially disturbing spatial cues already present in the stereo recording.

If the actual binaural eardrum signals remained identical under all those conditions, then yes, I would expect them to sound identical. My question is whether a single static GRAS magnitude-FR measurement is sufficient to establish that they actually do.
 
Last edited:
The E3 is very position-dependent. Moving it just a few millimeters forward or backward on my head can significantly change the response above 4 kHz. For example, play a 5 kHz test tone and move the cups slightly - you can get anything from a deep null to a peak depending on the position and getting both ears positioned exactly the same is extremely difficult.

The E3 can sound great, but the positional variation and seal sensitivity make it less plug-and-play than I would like.
Yes, that's essentially my point. If the sound field over the pinna is strongly non-uniform and frequency-dependent, then even a small change in fit can change which parts of the pinna are being excited and therefore change the resulting transfer function to the eardrum.

In that case, a headphone could show a very good eardrum FR at one fixed measurement position, while being less stable on a real ear across small variations in fit and pinna geometry.

That could make the spatial cues contained in a stereo recording less consistently reproduced. This is at least consistent with my simulation, where the AMTS-type geometry showed substantially greater spatial variation over the pinna, particularly in the 5–10 kHz region.

I'm not claiming that this proves worse spatial perception. I'm proposing it as a possible mechanism that would explain the observation.
 
Why does the E3 sound spatially narrow despite measuring so well?

The E3 has extremely low distortion, an excellent frequency response, clean group delay, and apparently very good time-domain behavior. Yet both my 10-year-old son and I had the same impression with Yosi Horikawa's Bubbles: many of the balls seemed to fall into almost the same spatial location. I have also seen the E3 described as having a “giant IEM” type of presentation.

That seemed difficult to reconcile with the conventional measurements, so I built a simple 3D FDTD simulation, using Claude Fable 5 extensively to help with the implementation and analysis, to look at the sound field across the pinna rather than only at the ear-canal response. The models are deliberately simplified, and the AMTS geometry is approximated from photographs, so this is not intended as an accurate simulation of the commercial headphones.

Figure 1 — Simplified geometries and probe positions
View attachment 552264


Figure 2 — Incident level distribution across the pinna
View attachment 552265

The interesting part is that simple level variation does not seem to explain the difference. The DX10000CL-type model also has substantial spatial level variation.
When the overall level difference at each probe is removed and only the spectral shape is compared, the result becomes much clearer.


Figure 3 — All probe frequency responses. Top: including level differences. Bottom: spectral shape only.

View attachment 552266

Mean spatial variation of spectral shape from 5–10 kHz:
  • MDR-Z1R-type: 0.8 dB
  • LCD-type: 1.5 dB
  • DX10000CL-type: 1.3 dB
  • DCA/AMTS-like: 2.7 dB
I also separated smooth spatial gradients from local irregularity using a local planar fit.

Figure 4 — Local spatial gradient vs. residual irregularity
View attachment 552267


The AMTS-like model is a clear outlier in the 5–10 kHz spectral-shape result.
So my current hypothesis is simply this:

A headphone can have an excellent frequency response at the ear canal while still producing a highly position-dependent spectral field across the pinna.

If so, conventional single-point measurements may be missing a variable relevant to spatial perception. The “giant IEM” impression would also make some intuitive sense: the final FR can be excellent, while the pinna is no longer being illuminated by a spatially consistent wavefront.

This obviously does not prove that AMTS causes the E3's spatial presentation, but I think it is an interesting result worth testing experimentally.

In my expperience, the better a headphone tunnig tracks the harman target from 1k to 3k the more it feels like you have speakers right next to your ears.

Hifiman tunes to have a depression around 2k, supposedly because they are replicating a room Dr Bian likes to listen to music at.

In the end its all psycho-accoustics, your brain interprets that there is more space because its trained to associate that a bigger room attenuates those frequencies. But in reality the distance between the transducer and your eardrum changes very little from headphone to headphone, so sources of sound cannot be further away. And the enclossure that the headphone makes when its placed on your head just doesnt have sufficient space to let resonances happen and interfere with the main wave, the sound spatial queues we get from the environment come from those interferences.

But then again, my thoughts come from purely inductive thinking, someone might have actual research on the matter.
 
I agree with the limiting case: if two headphones produced exactly the same pressure waveform at both eardrums, including magnitude and temporal/phase behaviour, then the different physical mechanism used to produce it would be irrelevant.

But that is not quite the hypothesis I am testing.

In my simplified simulation, the AMTS-type geometry produces a much less spatially uniform near-field over the pinna, especially in roughly the 5–10 kHz region. More importantly, the pattern changes substantially with frequency rather than behaving like a simple smooth level gradient.

I am not claiming that the auditory system somehow “reads” pressure at different points on the pinna directly. The question is whether a highly non-uniform, frequency-dependent incident field makes the resulting headphone-to-eardrum transfer function more sensitive to the listener's actual pinna geometry and to small changes in fit or position.

If so, two headphones could measure similarly at the eardrum of one fixed GRAS configuration while behaving less similarly across real human ears or small positional changes.

And that distinction matters because the pinna-related spectral filtering used in spatial perception occurs mainly in this high-frequency region.

So the hypothesis is not “the headphone creates elevation information from stereo.” It is that a spatially non-uniform incident field may interact with the pinna less robustly, potentially disturbing spatial cues already present in the stereo recording.

If the actual binaural eardrum signals remained identical under all those conditions, then yes, I would expect them to sound identical. My question is whether a single static GRAS magnitude-FR measurement is sufficient to establish that they actually do.
Yes, I think your points are sensible now that you've expanded on the mechanism. And this is also what @Chagall with his position dependant E3 headphone. You'll see this position dependance of a particular headphone on measurement rigs like the one Amir & Oratory use, so you can still get useful information from that. Oratory in particular measures the headphone at different positions on the ear and the measurement he shows is the average of those, and that's what his EQ is based on. In my own case with wearing (& also measuring headphones on my miniDSP EARS), I find naturally central measurements to be of the most use though, because the headphone position on my head is accurate and mostly central (depending on headphone) each time I wear it so I don't think I get much variation. There has been some work done on how much headphones react & measure differently when worn on different real subjects using blocked ear canal measurements, and open headphones and particularly the HD800 & K702 for instance showed the least variation between subjects, so there is something to be said for choosing a headphone that is reliable across different human anatomies and slightly different positioning. I'll see if I can find the graph I'm referring to.....damn it, can't find it on my PC - it's a pic showing variation between different headphones from on head measurements, I think done by Harman.

EDIT: found the graph:
 

Attachments

  • On head headphone variation.png
    On head headphone variation.png
    452.6 KB · Views: 42
Last edited:
Yes, I think your points are sensible now that you've expanded on the mechanism. And this is also what @Chagall with his position dependant E3 headphone. You'll see this position dependance of a particular headphone on measurement rigs like the one Amir & Oratory use, so you can still get useful information from that. Oratory in particular measures the headphone at different positions on the ear and the measurement he shows is the average of those, and that's what his EQ is based on. In my own case with wearing (& also measuring headphones on my miniDSP EARS), I find naturally central measurements to be of the most use though, because the headphone position on my head is accurate and mostly central (depending on headphone) each time I wear it so I don't think I get much variation. There has been some work done on how much headphones react & measure differently when worn on different real subjects using blocked ear canal measurements, and open headphones and particularly the HD800 & K702 for instance showed the least variation between subjects, so there is something to be said for choosing a headphone that is reliable across different human anatomies and slightly different positioning. I'll see if I can find the graph I'm referring to.....damn it, can't find it on my PC - it's a pic showing variation between different headphones from on head measurements, I think done by Harman.

EDIT: found the graph:

Yes, and this is the main problem with the E3, and also in my humble opinion, a flaw in the way headphone measurements are currently presented. Some headphones vary in FR only across different heads, while the E3 seems to vary across heads with the possible seal issues in the bass, and additionally varies with position on the same head in the treble.

We do get useful information from measurement rigs, but we rarely get to see the positional variation itself. That is another topic altogether, but I think it is important information to have.

However, when positioned correctly, E3 sounds subjectively very good to me compared with my HD800S and HD 480 Pro, and that is probably the most frustrating part :)
 
Yes, and this is the main problem with the E3, and also in my humble opinion, a flaw in the way headphone measurements are currently presented. Some headphones vary in FR only across different heads, while the E3 seems to vary across heads with the possible seal issues in the bass, and additionally varies with position on the same head in the treble.

We do get useful information from measurement rigs, but we rarely get to see the positional variation itself. That is another topic altogether, but I think it is important information to have.

However, when positioned correctly, E3 sounds subjectively very good to me compared with my HD800S and HD 480 Pro, and that is probably the most frustrating part :)
Yep, and that graph I showed in my last post also shows the variation from seating to seating for the same head. You can see the HD800 is broadly just one line per person, so not really much variation from seating to seating, whereas some of the other headphones in that pic show variance between different seatings on the same person. HD800 is a bit of a win here inasmuch that it's not showing much variation between seatings & also not much variation between people. Also, K701/K702 looks to be performing just as good in both those terms as well.

I wonder how important either or both of those variables are in getting a good & predictable sound from your headphone. I think for sure you don't want much positional variation on your own head as it makes it easier to wear the headphone and get the same or very similar experience. The second variable of different blocked ear canal measurements for different heads, maybe that's just a natural function of that person's HRTF and it's just natural & similar to how the measurements would be if they were listening to speakers instead, but then you have the HD800 & K701/K702 that don't seem to show much variation from person to person - what does that say, does that say that the person's "pinna HRTF's" are not being incorporated properly into the final sound or does it instead say that the headphone gives a similar experience of sound to each different person & hence it's a more predictable headphone in terms of quality. I think on that last point it's easy to conclude that you don't want the bass falling off a cliff or being highly variable between different people because that's just showing seal issues, but above that I'm not sure.

I think it could be useful see positional information in reviews, in terms of how much variance you get with different headphone positioning, as that's a pretty strong case for a good quality aspect in a headphone.
 
After reading so much high praise for and even rapturous descriptions of the E3, the slide into "the problem with E3" mode and the propagation of excruciating fatal-flaw discourse seems like a chronic syndrome is our little world of perfectionist audio. A luta continua.


maxime.jpeg
 
In my expperience, the better a headphone tunnig tracks the harman target from 1k to 3k the more it feels like you have speakers right next to your ears.

Hifiman tunes to have a depression around 2k, supposedly because they are replicating a room Dr Bian likes to listen to music at.

In the end its all psycho-accoustics, your brain interprets that there is more space because its trained to associate that a bigger room attenuates those frequencies. But in reality the distance between the transducer and your eardrum changes very little from headphone to headphone, so sources of sound cannot be further away. And the enclossure that the headphone makes when its placed on your head just doesnt have sufficient space to let resonances happen and interfere with the main wave, the sound spatial queues we get from the environment come from those interferences.

But then again, my thoughts come from purely inductive thinking, someone might have actual research on the matter.
I understand that impression. When the presence region is elevated, it can definitely make the sound feel more immediate, as if it is playing right next to your ears.

But that is a somewhat different kind of spatial perception from what I am talking about here. For example, imagine a bouncing ball moving from the front-right, tracing an arc at a certain height, and then passing behind you to the left. I am talking about the ability to perceive that kind of 3D trajectory.

Depending on the recording, some of those spatial cues can already be encoded in the signal. In particular, the higher-frequency region, roughly around 4–12 kHz, is important for elevation and front/back perception because the pinna creates direction-dependent spectral peaks and notches through reflection, diffraction and interference.

My hypothesis about the E3 is that, in this high-frequency region, the sound pressure field over the pinna may be spatially non-uniform, and that the pattern of this non-uniformity may itself change with frequency.
If that is the case, then for listeners who rely strongly on pinna-related spectral cues, the E3 may reproduce those cues less consistently, making 3D spatial perception more difficult.

So I agree that 1–3 kHz tuning can strongly affect the subjective sense of distance or immediacy. I just think that is a different mechanism from the one I am trying to describe here.
 
Yes, I think your points are sensible now that you've expanded on the mechanism. And this is also what @Chagall with his position dependant E3 headphone. You'll see this position dependance of a particular headphone on measurement rigs like the one Amir & Oratory use, so you can still get useful information from that. Oratory in particular measures the headphone at different positions on the ear and the measurement he shows is the average of those, and that's what his EQ is based on. In my own case with wearing (& also measuring headphones on my miniDSP EARS), I find naturally central measurements to be of the most use though, because the headphone position on my head is accurate and mostly central (depending on headphone) each time I wear it so I don't think I get much variation. There has been some work done on how much headphones react & measure differently when worn on different real subjects using blocked ear canal measurements, and open headphones and particularly the HD800 & K702 for instance showed the least variation between subjects, so there is something to be said for choosing a headphone that is reliable across different human anatomies and slightly different positioning. I'll see if I can find the graph I'm referring to.....damn it, can't find it on my PC - it's a pic showing variation between different headphones from on head measurements, I think done by Harman.

EDIT: found the graph:
Yep, and that graph I showed in my last post also shows the variation from seating to seating for the same head. You can see the HD800 is broadly just one line per person, so not really much variation from seating to seating, whereas some of the other headphones in that pic show variance between different seatings on the same person. HD800 is a bit of a win here inasmuch that it's not showing much variation between seatings & also not much variation between people. Also, K701/K702 looks to be performing just as good in both those terms as well.

I wonder how important either or both of those variables are in getting a good & predictable sound from your headphone. I think for sure you don't want much positional variation on your own head as it makes it easier to wear the headphone and get the same or very similar experience. The second variable of different blocked ear canal measurements for different heads, maybe that's just a natural function of that person's HRTF and it's just natural & similar to how the measurements would be if they were listening to speakers instead, but then you have the HD800 & K701/K702 that don't seem to show much variation from person to person - what does that say, does that say that the person's "pinna HRTF's" are not being incorporated properly into the final sound or does it instead say that the headphone gives a similar experience of sound to each different person & hence it's a more predictable headphone in terms of quality. I think on that last point it's easy to conclude that you don't want the bass falling off a cliff or being highly variable between different people because that's just showing seal issues, but above that I'm not sure.

I think it could be useful see positional information in reviews, in terms of how much variance you get with different headphone positioning, as that's a pretty strong case for a good quality aspect in a headphone.

For example, imagine Taylor Swift singing directly in front of you while neither she nor your head is moving. In a natural sound field, the sound arriving at your pinna is consistent with a source coming from that direction, and your own pinna then applies its normal direction-dependent filtering.
If a headphone changes its high-frequency response substantially when it is moved only a few millimeters on the same head, then it cannot be reproducing that same pinna input in a very stable way.

In other words, it may reproduce the average eardrum response reasonably well at one particular position, but it is not necessarily reproducing the natural incident sound field that would exist if Taylor Swift were actually standing in front of you.
That is the part I think may matter for spatial perception. If the spectrum varies significantly across different parts of the pinna, and that spatial pattern also changes with frequency, then the pinna-related spectral cues reaching the ear canal may become less consistent.

So for listeners who rely strongly on those cues, this could be quite important. I would not expect the effect to be identical for everyone, since pinna geometry and sensitivity to spectral cues vary between listeners.
 
Yes, and this is the main problem with the E3, and also in my humble opinion, a flaw in the way headphone measurements are currently presented. Some headphones vary in FR only across different heads, while the E3 seems to vary across heads with the possible seal issues in the bass, and additionally varies with position on the same head in the treble.

We do get useful information from measurement rigs, but we rarely get to see the positional variation itself. That is another topic altogether, but I think it is important information to have.

However, when positioned correctly, E3 sounds subjectively very good to me compared with my HD800S and HD 480 Pro, and that is probably the most frustrating part :)
What sounds “natural” may actually vary quite a lot from person to person.

I don’t dislike the HD800S either, but my impression is that it can sometimes create an almost oversized Taylor Swift inside my head. The apparent size of the source feels a little larger than life, which can be slightly unnatural to me.
The MDR-Z1R, which I personally like a lot, gives me something much closer to a life-sized image, so it sounds more natural to me in that sense.
With the E3, on the other hand, I tend to get the impression of a very small Taylor Swift with a very small mouth.

So I suspect there may be quite a lot of individual variation in how people process pinna-related information, and in what apparent source size they perceive as natural. It is probably one of those areas where psychoacoustics gets very difficult.

That said, I still think one important factor in creating a natural spatial presentation is for the headphone to provide the pinna with a reasonably consistent and well-behaved incident sound field, rather than one with large spatial irregularities that change substantially with frequency or position.
 
I understand that impression. When the presence region is elevated, it can definitely make the sound feel more immediate, as if it is playing right next to your ears.

But that is a somewhat different kind of spatial perception from what I am talking about here. For example, imagine a bouncing ball moving from the front-right, tracing an arc at a certain height, and then passing behind you to the left. I am talking about the ability to perceive that kind of 3D trajectory.

Depending on the recording, some of those spatial cues can already be encoded in the signal. In particular, the higher-frequency region, roughly around 4–12 kHz, is important for elevation and front/back perception because the pinna creates direction-dependent spectral peaks and notches through reflection, diffraction and interference.

My hypothesis about the E3 is that, in this high-frequency region, the sound pressure field over the pinna may be spatially non-uniform, and that the pattern of this non-uniformity may itself change with frequency.
If that is the case, then for listeners who rely strongly on pinna-related spectral cues, the E3 may reproduce those cues less consistently, making 3D spatial perception more difficult.

So I agree that 1–3 kHz tuning can strongly affect the subjective sense of distance or immediacy. I just think that is a different mechanism from the one I am trying to describe here.

I get your theory, cannot comment on whether or not you are on the right track due to lack of knowledge.

I have some questions for the sake of completion.

The arc that the ball traces in your example, isnt identifying the trayectory with just 2 channels hard, if not imposible? I guess that if the mix was done to be listened on headphones or speakers (let keep it to 2 channels), and you were to listen on the other device spatial queues would be less effective. And then you are also dependant on crossfeed, at least on headphones, spatial location through sound seems to depend on using both ears simultaneously.

And then on the pressure map over the pinna, even if the freuquencies are shifted, the passing of the ball would sweep through the frequency range just the same, you should be able to tell that motion is being implied through sound regardless, or is there a specific reason this wouldnt happen?
 
I get your theory, cannot comment on whether or not you are on the right track due to lack of knowledge.

I have some questions for the sake of completion.

The arc that the ball traces in your example, isnt identifying the trayectory with just 2 channels hard, if not imposible? I guess that if the mix was done to be listened on headphones or speakers (let keep it to 2 channels), and you were to listen on the other device spatial queues would be less effective. And then you are also dependant on crossfeed, at least on headphones, spatial location through sound seems to depend on using both ears simultaneously.

And then on the pressure map over the pinna, even if the freuquencies are shifted, the passing of the ball would sweep through the frequency range just the same, you should be able to tell that motion is being implied through sound regardless, or is there a specific reason this wouldnt happen?
Basically, yes: with ordinary stereo playback and no individualized HRTF processing, the spatial image often tends to form inside the head rather than as a fully externalized 3D scene.
So in my bouncing-ball example, I am not necessarily saying that you would perceive a perfectly realistic ball moving through a real room from front-right to rear-left. It may instead feel more like a trajectory inside the internal sound field in your head.

My point is that the apparent depth, height and shape of that internal space can still depend on how the headphone interacts with the pinna.
This is also why IEMs can feel different. Since they largely bypass the outer pinna, they reduce the contribution of the listener's own pinna-related spectral filtering. You can still perceive left-right motion, timing differences, level differences and some depth cues from the recording, but elevation and front/back cues may be weaker or less natural for some listeners.

And regarding your last question: I don't think the issue is simply whether the ball's sound “sweeps through the frequency range.” The auditory system is not just tracking a moving frequency sweep. It is interpreting the detailed spectral shape at each moment as part of a spatial cue.
If the headphone-pinna interaction changes those peaks and notches in a way that is inconsistent with the encoded spatial information, the motion itself can still be audible, but the perceived 3D trajectory may become flatter, distorted or less stable.

So I would separate “I can hear that the sound is moving” from “I can reconstruct the intended 3D path of that movement.”

Incidentally, with a well-made binaural recording, it is also possible in some cases for the spatial image to externalize outside the head and feel much more like listening to real sources or loudspeakers in front of you.
 
For example, imagine Taylor Swift singing directly in front of you while neither she nor your head is moving. In a natural sound field, the sound arriving at your pinna is consistent with a source coming from that direction, and your own pinna then applies its normal direction-dependent filtering.
If a headphone changes its high-frequency response substantially when it is moved only a few millimeters on the same head, then it cannot be reproducing that same pinna input in a very stable way.

In other words, it may reproduce the average eardrum response reasonably well at one particular position, but it is not necessarily reproducing the natural incident sound field that would exist if Taylor Swift were actually standing in front of you.
That is the part I think may matter for spatial perception. If the spectrum varies significantly across different parts of the pinna, and that spatial pattern also changes with frequency, then the pinna-related spectral cues reaching the ear canal may become less consistent.

So for listeners who rely strongly on those cues, this could be quite important. I would not expect the effect to be identical for everyone, since pinna geometry and sensitivity to spectral cues vary between listeners.
And also if both drivers don't match each other in frequency response then you'll also get some imaging issues potentially, which again could come down to both if one earcup is positioned differently on your left ear vs your right ear, or it can result from measurable channel imbalance through the frequency range of the actual headphone hardware. So yes how you wear it (which is worsened if the headphone is very sensitive to positional changes) plus the inherent channel imbalance can affect the imaging you get. I think those general ideas make sense.
 
Back
Top Bottom