• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

“Human ears are better than measurements!” / “Humans can hear things we can’t even measure!”

Of course, stereo imaging and soundstaging have no objective reality, so can't be measured, only appreciated subjectively, as they're created in the brain of the listener. Room acoustics can be measured for things like RT60 and room modes, but again, the 'feel' of a room and the effect that has on the perceived sound is subjective, and in the mind of the listener. For electronics and even loudspeakers, there's nothing in their performance that can't be measured, although for any sighted listener, the cosmetic performance and knowledge of the manufacturer greatly affect the perceived sound.

S.


I mostly agree — the final “feel” of stereo imaging, soundstaging, and room acoustics is ultimately created and experienced in the listener’s brain, so it’s inherently subjective.

However, research shows it’s not purely unmeasurable. A 2025 study used non-verbal Multidimensional Scaling (listeners rated similarity between different real rooms without using words or scales) and identified clear perceptual dimensions. These correlated well with specific psychoacoustic parameters (echo density, roughness, fractal correlation dimension, etc.) — sometimes better than conventional metrics like RT60.

So while the experience is subjective, it is grounded in real, quantifiable acoustic features. The brain is responding to measurable things in sophisticated ways that standard measurements don’t always fully capture.

ST
 
A 2025 study used non-verbal Multidimensional Scaling (listeners rated similarity between different real rooms without using words or scales) and identified clear perceptual dimensions.
Do you have a link, that description really has me intrigued?
 
  • Like
Reactions: STC
Of course, stereo imaging and soundstaging have no objective reality, so can't be measured, only appreciated subjectively, as they're created in the brain of the listener.
My position is that if it can be heard, it can be measured. If we are willing to take the time to figure out how.

This includes soundstaging, imaging, anything at all. If it can be heard, then it is contained in the signal, and without it, the signal would be different.

The difficult part is finding the right measurement(s) to isolate the effect we are interested in.
 
My position is that if it can be heard, it can be measured. If we are willing to take the time to figure out how.

This includes soundstaging, imaging, anything at all. If it can be heard, then it is contained in the signal, and without it, the signal would be different.

The difficult part is finding the right measurement(s) to isolate the effect we are interested in.
Have you never miss heard something? Hearing is falliable.
 
My position is that if it can be heard, it can be measured. If we are willing to take the time to figure out how.

This includes soundstaging, imaging, anything at all. If it can be heard, then it is contained in the signal, and without it, the signal would be different.

The difficult part is finding the right measurement(s) to isolate the effect we are interested in.
The problem is that not every brain processes sound the same way and each room/speaker has similar acoustics.
Soundstage, imaging, height and depth of sound, PRAT and all the other vague descriptions are a 'brain dependent' aspect (imagination) and as such is dependent on much more 'input' than just the sound waves arriving at your head.
If you can figure it out and prove it in a reliable way... let us know.
It is a perception thing not a technical aspect so there will be no 'relevant measurement'.

For instance now listening to a K240DF which measures poorly, has absolutely no bass and colored mids with poor treble seen from a hi-fi standpoint. And while not all music sounds even decent on it some recordings sound really 'pleasant' (after getting used to the presentation for an hour).
Going back to my 'go-to headphone' is a massive 'improvement' but until I do the K240 DF is pleasant. I get the same effect with a Koss KSC35 b.t.w.

Brains ... pfft.
 
Last edited:
All waves arriving at a particular point can be described by amplitude and frequency over time (phase being a function of the latter), not just radio waves.
I would venture all you need is amplitude vs time. Or sampling wouldnt work.
 
I would venture all you need is amplitude vs time. Or sampling wouldnt work.
Fair. I struggle a bit to get this phrase where it needs to be so that some of our periodic interlocutors can understand that the (entire) properties of soundwaves are known and measurable, using familiar tools.
 
Fair. I struggle a bit to get this phrase where it needs to be so that some of our periodic interlocutors can understand that the (entire) properties of soundwaves are known and measurable, using familiar tools.
PS, it’s interesting that many newly arrived subjectivists don’t make more of another property of waves: direction. How did the objectivists gain a monopoly on directivity? My thesis is that directivity research removes an element of mystery taking out all the fun (for some).
 
The topic of level differences / just noticeable differences (JNDs) recently came up in another thread. I have looked at some literature again to potentially get better data than that provided in the OP. This lead me to the following references:


which quotes

[A] Yost , Nielsen (1985), Fundamentals of hearing: an introduction
https://openlibrary.org/books/OL2838789M/Fundamentals_of_hearing

which itself presents data taken from

[B*] Jesteadt et al. (1977), Intensity discrimination as a function of frequency and sensation level
https://pubs.aip.org/asa/jasa/artic...896/Intensity-discrimination-as-a-function-of

And I also found these two additional relevant publications:

[C] Johnson et al. (1993), Just noticeable differences for intensity and their relation to loudness
https://pubs.aip.org/asa/jasa/artic...Just-noticeable-differences-for-intensity-and

[D] Scott et al. (2003), The Subjective Loudness of Typical Program Material
https://aes2.org/publications/elibrary-page/?id=12350

[A] and [B*] show us that the JND for pure tones in decreases with SPL: While it is around 1-1.5 dB at 20 dB, it scales down to 0.4-0.6 dB at 80 dB with the slope indicating that we might get closer to 0.2 dB JND around 100 dB_SPL:
Yost Nielsen 1985.png

But that extrapolation isn’t a given. There also appears to be a dependency on the stimulus frequency, but the order of frequencies is all over the place and there are no error bars in the plot which would be necessary to determine if the differences are significant and how reliable the data is. In the original study [B*], the range of 600 - 1000 Hz showed the smallest standard deviation for the experiments, but not necessarly the lowest JND. The latter occured in that range for 40 and 80 dB_SPL, but landed at 200 Hz for 5 dB_SPL and 4000 Hz for 10 and 20 dB_SPL average levels.

In [C], Johnson et al. used pure tones with an optional masking noise (Wide Band Noise or Narrow Band Noise) to determine, among others, the JND. In this case, the minimum JND of 0.2 dB is reached at 60 dB_SPL with wide band noise:
Johnson 1993.PNG


And then there is Scott et al. [D], where different types of music and noise were investigated. Here, white noise was the most discriminating stimulus with an average matching error of 0.16 dB vs around 0.6 dB for various types of music:
Scott 2003_1.PNG

For some subjects, their best level matching score for noise reached around 0.08 dB, while the best score for music (so excluding the „Ad“) is 0.15 dB for one listener, with everybody else scoring above 0.2 dB for all types of music:
Scott 2003_2.PNG

The study by Scott et al. thereby seems to suggest that humans might be able to distinguish level differences below 0.1 dB for white noise and possibly as good as 0.15 dB for music. Sadly, there is no measurement of the threshold for pure tones in the paper which would allow for a direct comparison to JND in noise or music using the same methodology.

In addition, the methodology of the experiment in [D] has one significant flaw: It is not an ABX test. Instead, participants were presented with different attenuation levels of the same signal in relation to a reference and could then use a volume control with 0.1 dB steps to match them as best as they could. As far as I understand the test setup, this allows for a form of „cheating“: Even if you can only detect JND of +-0.5 dB reliably, you can still turn the volume up until you reach your personal smallest positive JND, then turn it down until you reach your smallest negative JND counting the volume steps. Stepping back exactly half the number of steps between those two thresholds will lead to a very precise volume matching result which will be much better than what you could achieve in a true blind ABX test. From my perspective, it can not be excluded that such a method was applied by some of the listening test subjects.

So as it stands now, the JND range listed in the OP with around 0.25 dB being commonly listed and some people under some circumstances reaching closer to 0.1 dB still seems plausible to me for music. I encourage anybody who’s interested in this to post sources which show a different outcome.
 
Last edited:
Out of interest and to test the boundaries I have done an ABX comparison using pink noise generated at -6 dB and then attenuated by 0.5 dB down to 0.05 dB. The results were:

Attenuation vs. referencefoobar2000 ABX scorep-value
-0.5 dB10/100.1%
-0.3 dB8/105.5%
-0.2 dB8/105.5%
-0.15 dB8/105.5%
-0.10 dB7/1017.2%
-0.05 dB7/10, 6/10 (repeated)17.2%, 37.7%

I'm pretty confident I could score higher on the -0.3 dB run, but my patience today is limited ;) For -0.2 dB and -0.15 dB, I'm sceptical I could repeatedly score 8/10. For -0.1 dB and below, I was essentially guessing but I'm surprised that my guesses did not lead to a single score below 6/10. I think we would need closer to 40 runs to get reliable results for those.

In conclusion, for critical listening tests I would not be happy with level matching of worse than 0.1 dB. Even considering that Scott et al. (2003) showed that noise is more revealing than music for SPL differences, you want to be on the safe side here.
 

Attachments

Out of interest and to test the boundaries I have done an ABX comparison using pink noise generated at -6 dB and then attenuated by 0.5 dB down to 0.05 dB. The results were:

Attenuation vs. referencefoobar2000 ABX scorep-value
-0.5 dB10/100.1%
-0.3 dB8/105.5%
-0.2 dB8/105.5%
-0.15 dB8/105.5%
-0.10 dB7/1017.2%
-0.05 dB7/10, 6/10 (repeated)17.2%, 37.7%

I'm pretty confident I could score higher on the -0.3 dB run, but my patience today is limited ;) For -0.2 dB and -0.15 dB, I'm sceptical I could repeatedly score 8/10. For -0.1 dB and below, I was essentially guessing but I'm surprised that my guesses did not lead to a single score below 6/10. I think we would need closer to 40 runs to get reliable results for those.

In conclusion, for critical listening tests I would not be happy with level matching of worse than 0.1 dB. Even considering that Scott et al. (2003) showed that noise is more revealing than music for SPL differences, you want to be on the safe side here.
Something to consider is the impact of the listening position itself. Just moving your head slightly will have an impact and may skew the results.

Looking at there results it seems that differences as small as 0.05 dB are potentially audible (even if it might take a lot of time to get to statistical significance). Even differences you can barely hear might still skew the results over time and gradually add up (especially if you do a long test like 40 trials.
 
This post is intended as a reference for when the same question comes up the 1000ᵗʰ time and you don’t feel like typing out why the answer is “NO” again.



Nope. Human hearing is vastly inferior to modern measurement instruments – in most dimensions by multiple orders of magnitude. It’s not even close. We can measure everything we can hear and much more than that on top.

This whole idea of ears somehow being superior is a non-starter, because the majority of the music we listen to and all the music audiophiles describe as “challenging” like classic music, jazz or female singers, has been captured using microphones and standard commercial analog-to-digital converters (ADCs) - which are measurement instruments. So the whole discussion could simply end here, because there is no counter-argument to this: If you think you can hear things which can’t be measured, then how the f*ck could those things ever be part of the recorded audio signal, which was generated by measuring it in the first place? As you can see, the rest of this post is technically redundant because the initial proposition has been debunked.

Let’s nonetheless look at how much better instruments are than ears, just to get a grasp of the amount of hubris which was required to come up with that idea. The following are some important quantities you could measure in an audio signal:
  • Level thresholds
    Humans hearing is limited to a range of about 0 dB SPL at 1 kHz (threshold of hearing) to 130±10 dB SPL (pain threshold) [1, 2]. Even consumer-grade instruments like a MiniDSP UMIK 2 reach about -30 dB SPL @ 1 kHZ in self-noise [3, 4]. Their upper limit is usually around 120-140 dB SPL, depending on the microphone type [5]. Technically, capturing even higher SPLs would be possible, but there are very few use cases where this would be required, which is probably why it’s uncommon for mics to support it.

  • Level differences
    Humans are actually pretty decent at hearing level differences as small as 0.25 dB for pure tones, with some people in some tests being sensitive down to about 0.1 dB [6, 7] (you can test it yourself here). However, there are multiple effects at play limiting that ideal case threshold during normal listening, like the hysteresis effect [8]. Measurement instruments are more precise than human hearing for level differences, with good analyzers reaching 0.03 dB [9]. But in practice, when measuring with microphones and outside of soundproof chambers, environmental factors like fluctuations in the room noise limit the achievable precision for humans and instruments alike.
    If you need to measure precise level differences on speakers or headphones, it is therefore essential to measure the driving voltage using a multimeter instead of measuring the SPL using a microphone or SPL meter. Even cheap multimeters measure down to 1 mV, giving you a much higher precision than any microphone (and also surpassing humans by a good margin).

  • Noise
    Since instruments can detect sounds at much lower levels than humans, the same is true for noise. As noise can usually be analyzed over longer time scales, whereas instantaneous values may be required for level differences, instruments can gain further sensitivity in detecting noise by increasing the analyzed time frame. Consequently, modern audio analyzers can detect noise at absurdly low levels (< -120 dBFS) [10, 11]. In an odd twist of fate, modern DACs have reached noise levels which are at or below those of the best currently available audio analyzers – but not better than the best lab-grade instruments like precision voltmeters.

  • Distortion
    Humans can hear distortion in real music down to about -40 dBFS (= 1%) with instant A/B switching [12]. For speakers, it is generally assumed that distortion below 0.3% or -50 dBFS won’t be audible [13]. However, with pure tones in very favorable (artificial) scenarios, some people can reach a threshold of close to -80 dBFS with very high order harmonics (>20th order) [14]. Even better consumer-grade ADCs like the E1DA Cosmos can easily detect distortion down at -140 dBFS, though, which is just insanely far away from anything any human could ever even dream of perceiving [15].
    The reason the threshold numbers for humans diverge so much is that THD and IMD are not well correlated to the audibility of distortion: Distortion components close to the original tone can be fully masked even at very high levels, while higher order distortion can be easier to detect due to the lack of masking [11, 16]. Nonetheless, absolute lower thresholds have been established for THD and IMD and are valid. Multiple psycho-acoustic models exist which correlate significantly better with the audibility of distortion [17]. But in practice, manufacturers and reviewers alike publish THD or IMD numbers: They are easy to measure with great precision and are the de facto industry standard.

  • Frequency thresholds
    Human hearing is usually accepted to be limited to a range of 20 Hz – 20 kHz for young listeners and the upper hearing threshold quickly degrades with age [18, 19]. But even if you assume that we might hear slightly above or below those limits, instruments can measure anything from 0 Hz to multiple 100 kHz for audio analyzers and into the MHz and even GHz range for oscilloscopes and spectrum analyzers [20, 21]. So, no chance for humans to compete.

  • Frequency differences
    Modern audio analyzers are precise to 10⁻⁵ Hz or better, whereas humans usually fail to detect differences below 1.0 Hz in test tones [10, 22]. There’s just no contest here.

  • Time precision (“Timing”)
    For better consumer ADCs capturing signals at 192 kHz and 24 bit, the intersample time is 5.2 µs. However, the actual time resolution is significantly higher due to the sampling theorem and clears the nanosecond range [23, 24]. In addition, different measurement instruments like consumer-grade oscilloscopes can easily display a fully resolved 50 MHz sine wave, which takes 20 ns for one full oscillation. This means the time resolution of the scope must be significantly better than that. For humans, the agreed upon number is somewhere between 7 and 18 µs, which is multiple orders of magnitude worse than instruments [25, 26].

  • Phase differences
    As phase differences are effectively time differences for specific frequencies, the same arguments concerning timing (and frequency) precision made above hold and human ears are far outclassed by measurement instruments.
This shows how much better instruments are in the most relevant dimensions of audio quality. But we are not quite finished.

Let’s also discuss some typical counterarguments that come up when talking about measurements:
  • But I can hear differences, even though the measurements say there are none!”
    There’s three options:
    • Your specific device(s) are defective and there actually are differences where there would be none in an undamaged device.
    • You did not account for all boundary conditions and tested incorrectly. For example you may have missed to level-match to within 0.2 dB (ideally 0.1 dB) or you are using audibly different filters, EQ or other settings. Your seating position may have changed, or you may have taken more than a couple of seconds to switch between the settings or devices you tried to compare.
    • You are mistaken and are a victim of bias, like any other human could be. Or in other words: Your brain is tricking you and the differences would instantly vanish in controlled (double) blind testing. This is by far the most common and prevalent error in audio testing when done by a layperson.

  • But there’s more to an audio signal than time and voltage/current!”
    No, there absolutely isn’t. You can watch everything which makes up the analog audio signal on a two-dimensional oscilloscope screen. There’s time, there’s amplitude, that’s it. As amplitude, you can either measure voltage or current. The current is a direct result of the voltage being present: If two audio signals have an identical voltage curve on a specific system, their current curve will also be identical.

  • Science hasn’t fully understood it, yet!” / “We discover new things all the time!”
    If that were true and even assuming that the science behind some basic electronics somehow were incorrect, you would still easily be able to prove that such a problem exists using well controlled double blind listening tests. As long as no such evidence exists, that whole assumption is flawed: You can’t postulate major errors in some of the most well understood fields of engineering without providing any evidence and expect people to take you seriously.
    The reality is that in relation to electric and acoustic signals (like analog and digital audio), science has understood these very well for over half a century. We – for the most part – also know what we don’t know. We have applied knowledge from all the fields of science and engineering related to these topic to build devices like acoustic metamaterials or phased array antennas. If there were any fundamental errors in the basic science behind it, these things would simply not work.
    The fact that an individual might not understand all the theories and calculations used to design a specific device does not invalidate the science behind it. On the contrary, if an opinion is uninformed, the opinion itself should be the first thing to be questioned.
With that being said, the fact that measurement instruments are superior to human hearing in every way doesn’t mean that we can perfectly interpret or link measured quantities to every subjective sensation a human might perceive. A couple of good examples for this usually come up with speakers: We can measure their frequency response and directivity and how they interact with a specific room. But from those measurements, it might still be difficult to tell if one specific speaker in one specific room might have a “forward” representation of voices or create a good representation of the “sound stage” encoded in your recording.

In a similar manner, old school tube amps often add copious amounts of distortion to the signal, which some listeners describe as pleasing up to a certain degree. We can precisely measure the distortion, but we might not be able to predict today what exact amount is perceived as pleasing and at what point it becomes annoying again. But we can say when it would be audible or inaudible. So for subjective qualities where we can pinpoint specific measured quantities as their source or threshold, we absolutely can predict how and if those will be perceived by humans.

The fact that we don’t know all links between subjective impressions and objective data yet also does not mean that two devices which measure identical in every way might still sound different under otherwise identical boundary conditions: They wont. There’s no “more identical” past “identical”. The only reasonable caveat here is the question if we measured all relevant quantities for a specific device to declare it “identical” to another one. The required measurements for this judgment depend on the type of device (DAC, amp, speaker, etc.) and the boundary conditions under which it is expected to perform. One number alone certainly won't cut it, but nobody around here argues that it will.


Note: If you happen to have better sources for some of the above bullet points, feel free to share them below. If I made a mistake, please let us know how and where.
I sent this to some "high-end" audiophiles and let's just say, they lost their minds and all their marbles!
 
Something to consider is the impact of the listening position itself. Just moving your head slightly will have an impact and may skew the results.

Looking at there results it seems that differences as small as 0.05 dB are potentially audible (even if it might take a lot of time to get to statistical significance).

Or not get to it.

It's not a given that results like these will 'get better' with more trials.

(The # of trials per test should be determined in advance, to avoid 'cherry picking'....i.e., stopping when they 'look better')


Even differences you can barely hear might still skew the results over time and gradually add up (especially if you do a long test like 40 trials.
Or not, if you aren't really 'barely hearing' them.
 
This doesn't diminish the value of objective measurements - they remain indispensable. It simply acknowledges that, at least for now, carefully controlled listening tests are still the ultimate benchmark for assessing how enjoyable a product is in real-world use.
Where are these carefully controlled listening tests? And how on earth do you propose we measure enjoyment, the DUT being the 100% isolated variable factor?

More scientifically controlled listening tests would be highly welcome.

Imagine if wine tasting was done this way though….
 
Where are these carefully controlled listening tests?
Some blind tests
compare lossless and lossy compression
This is the industry gold standard: BS.1116 : Methods for the subjective assessment of small impairments in audio systems

And how on earth do you propose we measure enjoyment, the DUT being the 100% isolated variable factor?
In two steps - first a proper ABX to find out if there is an audible difference. Then, if, and only if, there is, then another blind test based on "which one do you prefer?".
Imagine if wine tasting was done this way though….
It has been done, many times. Often people even fail to tell if a wine is white or red, once the tasting is blind.
 
Thanks but I’m asking where are the controlled listening tests in the mainstream equipment reviews published. It’s these reviews that continue to be a barrier between the subjective and scientific realms.
In two steps - first a proper ABX to find out if there is an audible difference. Then, if, and only if, there is, then another blind test based on "which one do you prefer?".
Thanks. I understand that’s the protocol for establishing preference, but how would you drill down further and test for enjoyment?
It has been done, many times. Often people even fail to tell if a wine is white or red, once the tasting is blind.
Exactly!
That’s one big reason why manufacturers and the audio testing world (their customers) are against promoting the scientific approach.
 
Thanks but I’m asking where are the controlled listening tests in the mainstream equipment reviews published. It’s these reviews that continue to be a barrier between the subjective and scientific realms.

Those aren't our problem. ASR has no control over them. Why are you even asking?

Thanks. I understand that’s the protocol for establishing preference, but how would you drill down further and test for enjoyment?

Again, it's called preference testing, and it has a long history in multiple consumer product areas, and it has been crucial to the advances in loudspeaker design since the 1980s.

Are you saying you can 'prefer' something that you 'enjoy' less?

The 'audio testing world' is not the 'customer' of manufacturers. 'Reviews' of audio gear in mainstream publication rarely do any credible testing, and where they do (Stereophile), they separate it from the 'review'.
 
Back
Top Bottom