• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

An alternative semi-unique approach on Headphone/IEM "reference" frequency response targets.

theniroxy

New Member
Joined
Jul 12, 2026
Messages
2
Likes
3
Note: This is going to be my first post on ASR. I'm open to any feedback, correction or response about what I am talking about. Please show me where I messed up.

When we talk about "reference" frequency response targets, the usual approach is usually a [DFHRTF response curve] + [-0.8 dB/oct to -1.2 dB/oct, usually -1 dB/oct] tilting to create a room-like response with objective HRTF measurement data. The current "meta" of the DFHRTF response curves on IEMs is JM-1 DF created by Joel Merrifield, that is based on ISO 11904-1 averaged DFHRTF measurements of different kinds of people. Targets based off of this HRTF response usually get great feedbacks. Since IEMs completely bypass the pinna curve in the ear and all the upper midrange/treble is produced and amplified by the IEM driver, you'll need a so-called "PopAvg-DF" which JM-1 solves; as a rival to eliminate HRTF mismatches for most of the people (the average people). Otherwise to people's ears, it ...won't sound well, rather sounds harsh. Headphones (OE & AE) on the other hand, does not bypass the pinna. The driver's upper mids/treble amplification is not that significant on those kind of devices compared to iems, since the pinna amplification is still significantly active. That's exactly why we use measurement rigs with artificial pinnae or HATS's such as GRAS 43AC-7, B&K 5128, HMS II.3 to mimic and standardize reactions of the human pinnae. And the usual HRTF response curves used for creating "reference" headphone frequency response targets are the measurement rigs' lab DF responses for measurements on that rig such as 5128-DF, KEMAR DF (either KB50xx or KB006x), HMS II.3 DF. These curves are producing a strong match of the measurement rigs' pinnae response for the measured headphone. And that will translate to a significant match to the listeners' HRTF also, since their pinnae is also amplifying actively. However, HpTF variations doesn't end there since what you hear is a unique interaction of your head, pinna, pad seal, seating, how much does the pad touch your pinna, how much the headphone amplifies the pinna gain region vs your ear (the headphone type).. and so on that we can't create a static reference of.

Putting these aside, my approach is a semi-unique approach that I've been seeing and studying for a ..fair amount of time, so let me share it.


a) A Minor Problem With Raw Diffuse Field Reference;

The diffuse field reference is widely used in headphone/iem targets as I've pointed about. The reasons that it's being used are that its capability of measurements and its closeness to what we hear in real life. However, it is kind of a utopia. DF measures that every sound wave at every frequency is indirect. That is why it is called "diffuse field". However; in an idealized room that we are referencing off of (usually ITU-R BS.1116-3 compliant), the sound waves are not completely indirect. Even if we use a tilted/adjusted diffuse field response, it may still lack room-like natural perception. So in my approach, I prefer to add some minor/"microscopic" filters on the upper mids region to the referenced DFHRTF curves, based off of the SoundGuys target. Feedbacks of this target on web provided me information about the pinna gain region of this target curve is more "natural" than the usual DF reference to people. This effect cannot be verified through measurements, it has to be tested but that is the closest thing that I have got. And since the SoundGuys target is a preference target that does not aim for being "reference" and neutral, I did not rely on that target curve fully.

graph1.png


b) My Objective Reference Points before adding any filters;

Like the current approach, I used JM-1 DF for IEMs and the "rig DF" for headphones.

c) Problem With Tilting The Graph and The Room Effect;

The usual "rule" for standardized in-room response for speakers (usually ITU-R BS.1116-3 compliant rooms) is "-1 dB/oct tilt" on an anechoically flat loudspeaker. However; this is just an estimation, there is usually a transition/schroeder frequency. When we use the -1 dB/oct logic to headphones or IEMs using an HRTF response curve, it's most likely going to become a little bit not room-like and may be perceived as "bloated" around 200hz since it's mimicing a perfect line. That's where, @Sean Olive 's (The Harman Targets) and also @crinacle 's approach on IEF Preference 2025 comes in. I used high shelf and low shelf filters to adjust my target. When we do that, the schroeder frequency is being mimiced well and we will get a more room-like sound.

graph2.png


d) My Approach On Neutral Adjustments;

d.1) For my target curve with this approach, a -5 dB high shelf filter at the corner frequency of 2500 hertz is the almost equivalent to the "10 dB tilt-neutral" and it is enough to mimic an in-room response of loudspeakers.

d.2) As for the (sub)bass shelf, the perceptual neutrality changes with a couple of factors such as perceived tactility boost, minor HpTF variations, ear canal impedance, age factor and basic perception of sound by well ..your brain. My approach on this is not making a single line, rather making a range. My baseline bass shelf starts at +3 dB (maybe +3.5) to match the "10dB tilt-neutral" since the amount of bass on this curve is considered neutral to a lot of people and also audiophiles. The bass frequency is the Harman Standard by @Sean Olive which is 105 hertz for most of my curves. Unlike the most of my curves, my IEM target curve on B&K 5128 uses an 80 hertz bass shelf, since it delivers more tactile sound in direct-ear-canal setup. Just like @crinacle 's approach on the IEF Preference 2025 curve on B&K 5128 HATS.

graph3.png


Also there is also what i call the "Honesty Threshold". This is the threshold that most of the people perceive the bass shelf as "neutral" and not bassy. Which is usually the Harman IE 2019 v2 Target's bass shelf (on the 5128 version) at IEMs, and Harman OE/AE 2015 target's bass shelf at headphones.

graph4.png

graph5.png


e) Problem With IEC 60318-4/711 couplers;

As well-known, these kind of couplers may overestimate 8 kHz region, is unreliable after 10 kilohertz, and also smooth and amplify the bass curve artificially. Whilst making the compensation (the adjustments i talked about may differ on the 711), some things may be inaccurate and could cause mismatch with some couplers' behaviour. Besides that, since the 711-compensated verison of JM-1 is a delta from B&K 5128, the 500 hertz "bump" you see is a compensation for the B&K 5128's "rocking mode" problems with some IEMs. That's why in my compensation, i also adjusted 500 hertz on 711.

graph6.png


f) Silly Name;

I made the name of my target... the... *ba dumm tss*.. LIAR Target. It is a memorable and a kind of contradicting acronym. LIAR stands for Loudspeakers In A Room.

If you want to see my target curve based on my approach, the TXT files are attached.
Please share your feedbacks, I need enlightenment.

"Don't be hard on me, this is just my theory, new there, may not know things that deeply :,("
 

Attachments

Last edited:
The current "meta" of the DFHRTF response curves on IEMs is JM-1 DF created by Joel Merrifield, that is based on ISO 11904-1 averaged DFHRTF measurements of different kinds of people. Targets based off of this HRTF response usually get great feedbacks.
That and $5 will get you a cup of coffee.

The only way to verify validity of a target is to present a controlled preference studies with statistical analysis and comparison to other targets. As Sean Olive performs. No amount of words, references to this and that suffices. Such testing has shown DF to be a poor target to say nothing of the fact that a) true DF doesn't exist and b) no room has this kind of response.
 
Overall I think this is fine to experiment with but there are a few factual errors I'd like to correct:
However; in an idealized room that we are referencing off of (usually ITU-R BS.1116-3 compliant), the sound waves are not completely indirect. Even if we use a tilted/adjusted diffuse field response, it may still lack room-like natural perception. So in my approach, I prefer to add some minor/"microscopic" filters on the upper mids region to the referenced DFHRTF curves, based off of the SoundGuys target.
So, the reason we use Diffuse Field is not because the speaker-in-room setups used for making music are best characterized by a Diffuse sound field. It's to do with the brain bringing totally different processing to the table with headphones vs. speakers in rooms.

In the real world, our perception and identification of sounds around us goes through three steps:

1) The sound arrives at our eardrums, already filtered/affected by the frequency response (HRTF) effects that our ears commit to the incoming sound based on that sound source's location

2) The brain—which has had our entire life to condition itself to know both how our ears affect incoming sound, and the exact effects that different sound source locations commit to the sound—uses its knowledge to contextualize the sound that arrives at the eardrum

3) Our perception receives "the sound," but only after the brain has taken the parts its familiar with: frequency response cues it knows to associate with the effects of our anatomy (both generally, as well as specific to certain sound source incidences) are subtracted from what we actually perceive.

This is yes, a consequence of simple habituation, but the "incidence-specific" frequency response cues are also used to help know where sound sources are located around us based on sound alone, along with other aspects of HRTF like ITD/ILD. The important thing here is that the typical effects our ears commit to all incoming sound regardless of direction get subtracted from our perception, but the typical effects our ears commit to incoming sound from a specific direction get subtracted as part of the same processing to identify the sounds around us.

Sound arriving through headphones or IEMs is still being processed by our brain, which is trying to use everything it has in the toolkit to contextualize what it's hearing to make it sound normal... but it is simply missing a big chunk of context: the sound isn't playing from anywhere around you, it's playing from a fixed location on your head, following your head wherever it goes. This means the part of the processing concerned with "identifying/categorizing frequency response cues as indications of sound source location" is missing basically all of the information it would need to properly contextualize these frequency response cues.

So with headphones, the brain—like with sounds in the real world—still uses our "average" HRTF as the baseline expectation for "this is generally how our ears effect incoming sound," but because the brain doesn't have any of the cues provided by a sound source at a distance (no ITD/ILD, no crossfeed, no change in frequency response when the head moves), it can't actually apply all of the same processing on ingest the way it does with sounds in the real world. It only has enough information to apply the "average" HRTF processing.

So while the brain only applies the ingest processing for the "average" HRTF effects, any other frequency response features, even if they do resemble certain HRTF features that would arise due to source incidence, are no longer being contextualized by the processing that would make them sound normal, so we hear deviations from the "average" HRTF as colorations, not as indications of sound source incidence.

How are my headphones supposed to know that they should be applying eg. my 30 deg 0 elevation frontal HRTF as the timbral baseline to the incoming sound when there's no sound source in front of me... yk? This is just to say: Diffuse Field is not trying to approximate the sound of speakers in a room, headphones are not sound sources in rooms, so we have to treat them differently.

Separately from that, I would advise using the Headphones.com IEM Diffuse Field HRTF as your pre-adjustment baseline, if only because it actually includes the parts of the rig that do interact with the IEM (the ear canal transfer function of the 5128). I like your idea of shifting the primary resonance around 3 kHz lower in frequency (this would be necessary for those whose ear canal is longer than 5128 to get a proper match in primary resonance frequency), but I've not exactly found an ideal way to do this with actual objective/mathematical certainty.
The usual "rule" for standardized in-room response for speakers (usually ITU-R BS.1116-3 compliant rooms) is "-1 dB/oct tilt" on an anechoically flat loudspeaker. However; this is just an estimation, there is usually a transition/schroeder frequency. When we use the -1 dB/oct logic to headphones or IEMs using an HRTF response curve, it's most likely going to become a little bit not room-like and may be perceived as "bloated" around 200hz since it's mimicing a perfect line. That's where, @Sean Olive 's (The Harman Targets) and also @crinacle 's approach on IEF Preference 2025 comes in. I used high shelf and low shelf filters to adjust my target. When we do that, the schroeder frequency is being mimiced well and we will get a more room-like sound.
I'm not sure I follow here; the tilt is a fairly solid approximation of what'll happen when you bring a typical good (anechoically flat) speaker into a room, but to see what this would actually look like objectively, its worth looking at the actual ERDI/SPDI curves of speakers. What you'll see is that many actually exhibit the opposite of the behavior you mention:
1783909322286.png

Data courtesy of Amir, you'll see that the effects of indirect sound (what would start to downslope the speaker's response when put in a room) are more like a gradual negative high shelf filter placed somewhere around 500 Hz—most of the slope's magnitude shift is reached by 1 kHz, and then it flattens out. For this reason, this speaker (and honestly, a lot of good speakers that don't do fun tricks with directionality in the bass) will actually end up looking closer to a flat tilt—or even just a low shelf/low-order high cut filter that stops cutting around 200 Hz—when placed in a decent room than Harman's adjustments (105 Hz LS or 2500 Hz HS). I have a post getting into this more here, would love to know your thoughts on it.
As well-known, these kind of couplers may overestimate 8 kHz region, is unreliable after 10 kilohertz, and also smooth and amplify the bass curve artificially. Whilst making the compensation (the adjustments i talked about may differ on the 711), some things may be inaccurate and could cause mismatch with some couplers' behaviour.
So actually, the 711 coupler is potentially closer to accurate in the "8 kHz" region (put in quotes because its less damped at the half-wave length mode resonance, which will shift based on sound source distance from the microphone) than the 5128 is! Yes, 711 is still less accurate both above and below this length mode point, but that one point you mention is actually where it's arguably most accurate to humans if we're talking "human-like" insertions (ie. not at the reference plane)
Besides that, since the 711-compensated verison of JM-1 is a delta from B&K 5128, the 500 hertz "bump" you see is a compensation for the B&K 5128's "rocking mode" problems with some IEMs. That's why in my compensation, i also adjusted 500 hertz on 711.
Actually, the 500 Hz "bump" is a minor artifact that isn't to do with the rocking mode, it's an artifact of unknown source—if the IEM is stabilized in the ear canal and prevented from rocking, this artifact remains. I think it's probably fine to leave it in the delta compensation (the 711 version of the DF targets), as the disconnect observed in human measurements really isn't that large (pic related below, from Estimating the Sound Pressure at the Eardrum and Ear Canal Transfer Function with In-Ear Headphones by Etienne Rivet, Sean Olive, et al.) but by that same token, if you want to remove it I suppose it's likely not disastrous or anything.
.
Screenshot 2026-07-12 at 10.29.55 PM.jpg

Regardless, I think it's a very good idea to experiment with your own sorts of adjustments—EQ is an infinitely configurable tool, and we shouldn't stop at two preference adjustments if we think we can make it sound even better!
 
Last edited:
Welcome to ASR and for what it's worth, I appreciate the high effort that went into the first post. I don't have anything to add to the fairly sophisticated replies you've had already, but I do find the discussion of targets pretty interesting as it seems like a nearly intractable problem, even if you do have personal HRTF measurements!
 
Thank you for your insightful reply and corrections, @listener650 . As I didn't mention, I am not verifying any validity of any target. This is just an experiment based on an approach that I have made that is open to any discussion.

I'm not sure I follow here; the tilt is a fairly solid approximation of what'll happen when you bring a typical good (anechoically flat) speaker into a room, but to see what this would actually look like objectively, its worth looking at the actual ERDI/SPDI curves of speakers. What you'll see is that many actually exhibit the opposite of the behavior you mention
I must say that you're correct and thank you for the correction that I have made about loudspeaker behavior. It is completely valid to not follow.

However, as for my approach for headphone/iem translation, I'd rather stick to the Dr. Sean Olive's approach than adding a tilt to a DF curve, since it's perceived more "sweeter" translation, also Harman's preference tests statistically correlates with this as I think. I am not sure that if it could be classified as "reference", "neutral" or "uncolored" but honestly to my ears it sounds like a really well translation subjectively. I don't have the muscle to do any test about that, that's why I kept it open to replies.

So, the reason we use Diffuse Field is not because the speaker-in-room setups used for making music are best characterized by a Diffuse sound field. It's to do with the brain bringing totally different processing to the table with headphones vs. speakers in rooms.

In the real world, our perception and identification of sounds around us goes through three steps:

1) The sound arrives at our eardrums, already filtered/affected by the frequency response (HRTF) effects that our ears commit to the incoming sound based on that sound source's location

2) The brain—which has had our entire life to condition itself to know both how our ears affect incoming sound, and the exact effects that different sound source locations commit to the sound—uses its knowledge to contextualize the sound that arrives at the eardrum

3) Our perception receives "the sound," but only after the brain has taken the parts its familiar with: frequency response cues it knows to associate with the effects of our anatomy (both generally, as well as specific to certain sound source incidences) are subtracted from what we actually perceive.

This is yes, a consequence of simple habituation, but the "incidence-specific" frequency response cues are also used to help know where sound sources are located around us based on sound alone, along with other aspects of HRTF like ITD/ILD. The important thing here is that the typical effects our ears commit to all incoming sound regardless of direction get subtracted from our perception, but the typical effects our ears commit to incoming sound from a specific direction get subtracted as part of the same processing to identify the sounds around us.

Sound arriving through headphones or IEMs is still being processed by our brain, which is trying to use everything it has in the toolkit to contextualize what it's hearing to make it sound normal... but it is simply missing a big chunk of context: the sound isn't playing from anywhere around you, it's playing from a fixed location on your head, following your head wherever it goes. This means the part of the processing concerned with "identifying/categorizing frequency response cues as indications of sound source location" is missing basically all of the information it would need to properly contextualize these frequency response cues.

So with headphones, the brain—like with sounds in the real world—still uses our "average" HRTF as the baseline expectation for "this is generally how our ears effect incoming sound," but because the brain doesn't have any of the cues provided by a sound source at a distance (no ITD/ILD, no crossfeed, no change in frequency response when the head moves), it can't actually apply all of the same processing on ingest the way it does with sounds in the real world. It only has enough information to apply the "average" HRTF processing.

So while the brain only applies the ingest processing for the "average" HRTF effects, any other frequency response features, even if they do resemble certain HRTF features that would arise due to source incidence, are no longer being contextualized by the processing that would make them sound normal, so we hear deviations from the "average" HRTF as colorations, not as indications of sound source incidence.

How are my headphones supposed to know that they should be applying eg. my 30 deg 0 elevation frontal HRTF as the timbral baseline to the incoming sound when there's no sound source in front of me... yk? This is just to say: Diffuse Field is not trying to approximate the sound of speakers in a room, headphones are not sound sources in rooms, so we have to treat them differently.

Separately from that, I would advise using the Headphones.com IEM Diffuse Field HRTF as your pre-adjustment baseline, if only because it actually includes the parts of the rig that do interact with the IEM (the ear canal transfer function of the 5128). I like your idea of shifting the primary resonance around 3 kHz lower in frequency (this would be necessary for those whose ear canal is longer than 5128 to get a proper match in primary resonance frequency), but I've not exactly found an ideal way to do this with actual objective/mathematical certainty.
Again, thank you for your correction, you're correct. There is no angle hints on IEMs and headphones like on the direct sound of a speaker. Since we're eliminating the direct sound, DFHRTF measurements become the correct way. The "microscopic filters" I made is mostly based of my assumptions and some reviews that I'm also not sure myself, but as you said, it's definitely more open to be perceived as "colorations" rather than corrections. It might need a correction on the target itself that I've made.
Actually, the 500 Hz "bump" is a minor artifact that isn't to do with the rocking mode, it's an artifact of unknown source—if the IEM is stabilized in the ear canal and prevented from rocking, this artifact remains. I think it's probably fine to leave it in the delta compensation (the 711 version of the DF targets), as the disconnect observed in human measurements really isn't that large (pic related below, from Estimating the Sound Pressure at the Eardrum and Ear Canal Transfer Function with In-Ear Headphones by Etienne Rivet, Sean Olive, et al.) but by that same token, if you want to remove it I suppose it's likely not disastrous or anything.
Hmm... thank you for clarifying this dilemma. This 400-800 hertz dilemma on B&K 5128 rigs is something that I have encountered in a lot of IEMs, *strictly neutral* IEMs also. The perceived lack/change on that region is minimal, the "dip" shows in the graph doesn't translate to reviews that often. That's why I read 5128 graphs kind of differently. The reason that I have removed the "500hz bump" on the 711 compensation of JM-1 is exactly that, it's for better readings of frequency response graphs on the 711 couplers.

Anyways, I'm still open to more replies, thank you so much for your insight about this. It could sound way better if you contribute so!
 
That and $5 will get you a cup of coffee.

The only way to verify validity of a target is to present a controlled preference studies with statistical analysis and comparison to other targets. As Sean Olive performs. No amount of words, references to this and that suffices. Such testing has shown DF to be a poor target to say nothing of the fact that a) true DF doesn't exist and b) no room has this kind of response.
Sharur responded to this via a youtube video response haha
 
Back
Top Bottom