• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

The “Real” Speaker Preference Test

so I would either way expect products with excellent measurements, which are regularly winning blind/controlled listening tests, to be preferred under sighted conditions as well.

Stereophile, loudspeaker of 2019:

Notes on the Vote
Some POTY winners succeed by coaxing second- and third-place wins from the majority of our writers, but not this one: The Revel Performa F228Be won with three second-place votes and a full five first-place votes. The second-place Dutch & Dutch 8c—a product that I half-expected to take top honors—earned 4 points less and took second place


 
Harman's elevator muzac playlist could be sufficient for basic tonality tests, but it's limited.
That's very unfair, Harman was one of the very few who analyzed which "signals" enable the biggest discrimination of tonal problems (=resonances) at listeners and on first place was pink noise and close second the famous song from Tracy Chapman who has a very similar broad and dense spectrum and those are very far from "elevator muzac".

I sense a lot of criticism and bitterness from some towards the Harman research but I never see them propose any specific better alternatives except general text bubbles.
 
The topic of sales as a proxy for consumer preference has been studied. But it only tells you what consumers bought, not necessarily why they bought it or if they truly preferred it over other options. There are so many factors contributing to consumer hi-fi purchases, after all, and there's simply no possible way to isolate for actual listening preference. I haven't bought that much equipment in my life, but every time I've done so, availability, appearance, recommendations from others, brand admiration, and of course budget have mingled together to influence my decision.
So, if ´the science´ was successful in understanding people´s buying decisions, and leading to vastly superior products, we would see a lot of them now on the used market. Perfect timing. Where are they? Probably have never been produced and sold, would be a logical conclusion. Which coincides with what Dr. Olive was reporting, that Harman´s internal research institution turned towards earphones, headphones, immersive audio and surround, more or less abandoning loudspeaker research in the early 2010s.

But the used market is also an indicator of what people are dissatisfied with by today. Otherwise they would not sell it, right? While I can easily understand why people sell Focals or B&Ws in mass quantities - they have changed their sound tuning, they might not meet today´s expectation in terms of sound character, they are difficult to set up in a new room, majority of them is chunky - it is difficult to find an explanation for KEFs being sold. Other than lots of people are dissatisfied with the sound, which cannot be easily explained by sounding different in different rooms over moving. So, why?

Based on anecdotal data on the used market, I'm pushed to interpret that most people don't buy based on the science of good sound. They buy based on other factors, such as brand recognition, aesthetics, marketing and here is a big one: the lack of comparison in a controlled environment, you don't know what is good without bad, and vice versa. Keep in mind listening in a showroom with the salesman selecting the tracks and SPL for you, while other people are coming in and out of a showroom isn't the best experience to evaluate speaker sound quality.
 
Something like this is what MOST PEOPLE listen to...

1000001737.png
 
Since I spent a decent chunk of my career figuring out how to make people buy speakers, I can speak to this topic a bit more. My experience aligns with @Ellebob.

I surveyed a lot of consumers on how they decide to buy speakers or audio equipment in general.

If you ask people what they want in audio gear, the #1 answer by far is "Sound quality".

If you ask them to define sound quality, among the general public you don't get much further than "Bass" "Loud" "Clear sound". This is not because people can't make more detailed decisions ears-on, but the average listener doesn't have much audio vocabulary.

And, despite what people say about their selection process, it is impossible to overstate the impact of marketing. At least 6-10 years ago, it was incredibly common for "Bose" to be equated to good sound. In the minds of a shockingly large and diverse swath of the general public in the US, "Bose" was more or less the definition of good sound.

Measurements, at least for most people, are not a thing.

So, if someone actually gets as far as doing an in-person listening test, hyped bass, hyped treble, and simple loudness will win more often than perfect measurements. I mean, this is so well known and so common we have a name for it - "Showroom sound".

And, of course, aesthetics are a gating factor on basically any purchase, from $20 to $2M.

Basically, people really do want the best sound they can get within their budget and use case. OP's question is whether consumers actually prefer Harman sound. Their research shows they do. However, in practice, in the market, there are so many other factors at play that purely good sound (by the Harman standard or any other) is not nearly enough to achieve high sales numbers.
 
Last edited:
In addition, results in listening tests depend on program. Harman's elevator muzac playlist could be sufficient for basic tonality tests, but it's limited.

Choosing the program material for listening tests, and making sure all participants are sufficiently aware of how it was mixed/mastered, in my understanding is key to getting useful preference tests results. Mastering engineers who know a lot of different mixes, as well as recording/broadcast engineers judging their own, I would consider to be the most reliable when it comes to judging tonality and imaging.

I might not be aware of all playlists being used in the past, but particularly minimalistic recordings with close-mic´d (female) voices who without amplification ´would not fill an opera house´, as recording engineers like to put it, and added artificial reverb, are the least reliable to judge tonality. Natural reverb field is completely removed, and tonality of vocals depends vastly on microphone characteristics, placement and EQ.

Which explanations do you believe are stronger?

That was answered by @kimmosto already, but I would consider:

- choice of program material
- participants´ detailed awareness for material under reference conditions
- channel layout/stereo triangle geometry
- room acoustics, particularly bass

being most important. Note, that I am not saying these factors are leading to reliable or consistent results, but might explain why vast quantities of buyers at dealerships come to opposite preference decisions compared to those in controlled tests.

Harman was one of the very few who analyzed which "signals" enable the biggest discrimination of tonal problems (=resonances) at listeners and on first place was pink noise

That is an exact description of the circle of confusion going on with these tests. If you choose ´signals´ solely by the ability of listeners to identify slight tonal changes done by PEQ filters, you lay the ground for a discrimination test, not a preference test. For the latter, you need a reference ´in the real world´.

If the discrimination test thing would be taken more seriously, I suggest white noise, highpass-filtered @300Hz, and all tests executed with headphones. I am sure that this is even more ´revealing´ in terms of identifying narrow-banded resonance effects or bell filters than Tracy Chapman. But it is equally worthless when it comes to judging what sounds right, what sounds as intended by the recording engineer, and what delivers both tonal balance and reverb tonality.

I sense a lot of criticism and bitterness from some towards the Harman research

I don´t think this is aiming at Dr. Toole or Harman´s attempts in general. Rather the quite dogmatic way the results are treated by some as ´target curves that no-one must deviate from´ and how it is described as ´the one and only science´ in terms of sound quality.

His research indicates that the average listener will exhibit preferences for flat frequency response and low coloration in speakers IF AND WHEN ALL OTHER THINGS ARE EQUAL.

Not saying this very basic finding is wrong, I can more or less confirm it from own experience. Problem is, that for example taste regarding bass is showing quite some variations between participants in any tests, and it is basically impossible to isolate the frequency response from other properties when comparing existing loudspeakers, as no two loudspeakers have exactly identical directivity and decay behavior.

So people in a controlled blind test might statistically opt for the more (anechoically) linear model, but you have no clue if more exalted bass or more or less constant directivity were decisive factors as well. On the other hand, to identify the latter as decisive, you have to do further listening tests isolating one of the properties in question, which is not easy to do.

I have undergone such tests in an anechoic chamber with perfectly linearized anechoic response. Tonal differences solely resulting from resonance and decay were pretty pronounced.

Remember that I said many people get speakers home and then wonder why they don't sound anything like they did at the dealer's? An educated customer, aware of total speaker measurements ... including spinorama, distortion and frequency response ... will be better prepared to predict whether the speaker under consideration will sound acceptable in his home environment, and will do so long-term.

Fully support your theory and do think, this is a common problem in the real world.

The problem is, that commonly people who are told to ´solely rely on measurements´, are oftentimes as mislead as the ones in the stores being told to buy the speaker which sounds most exciting with a handful of demo tracks. Reason being anechoic response is mostly irrelevant when it comes to judging compatibility between room, speaker and desired placement, as direct sound FR can be easily corrected by DSP. Directivity index, response within the early reflection windows, response over the complete listening windows, geometry of bass sources potentially exciting room modes - these things are far more important in my understanding, and I see little effort from the ´measurement-objectivist´ side to really educate people here.

And I believe it is not true to say that the majority of buyers ignore or despise measurements. If that were true, ASR would not have the volume of traffic that it has.

I don´t think that traffic and purchases of rather expensive gear are anyhow connected. Many people come to articles on ASR by googling for products, many might simply like the fact that expensive stuff which they don´t have, occasions hefty criticism here. Admittingly, some might buy cheaper gear and find confirmation that this is well-measuring, which is absolutely legitimate, but not a hint that the market is really moved by belief in measurements.

The Revel Performa F228Be won with three second-place votes and a full five first-place votes.

Sounds like a jury of reviewers voted. I don´t really see the connection to enormous sales figures, or consumers opting for a particular speaker.
 
In the minds of a shockingly large and diverse swath of the general public in the US, "Bose" was more or less the definition of good sound.

Fully agree, but would go one step further: Bose was widely regarded as the definition of ´Better sound through research´, and exemplary for products representing ´the science´. After measurement-heavy marketing, specs records, and all the ´pseudo-physics marketing´ in the 1970s and 1980s, becoming increasingly irrelevant, Bose was the single most clever entity to take advantage of their reputation.

hyped bass, hyped treble, and simple loudness will win more often than perfect measurements.

As proven by JBL´s well-deserved success with portable speakers, party sound reinforcement and boomboxes. At least in these categories, the science on preference seemingly worked.
 
Last edited:
Fully agree, but would go one step further: Bose was widely regarded as the definition of ´Better sound through research´, and exemplary for products representing ´the scienc´. After measurement-heavy marketing, specs records, and all the ´pseudo-physics marketing´ in the 1970s and 1980s, becoming increasingly irrelevant, Bose was the single most clever entity to take advantage of their reputation.
Clever maybe, I don't know if outspending your competition by 10x on ads counts as clever though. ;)
As proven by JBL´s well-deserved success with portable speakers, party sound reinforcement and boomboxes. At least in these categories, the science on preference seemingly worked.
Yes, we were in that space, and they were one of two competitors that we felt were a threat in terms of sound quality... The other was House of Marley, believe it or not.
 
Since I spent a decent chunk of my career figuring out how to make people buy speakers, I can speak to this topic a bit more. My experience aligns with @Ellebob.

I surveyed a lot of consumers on how they decide to buy speakers or audio equipment in general.

If you ask people what they want in audio gear, the #1 answer by far is "Sound quality".

If you ask them to define sound quality, among the general public you don't get much further than "Bass" "Loud" "Clear sound". This is not because people can't make more detailed decisions ears-on, but the average listener doesn't have much audio vocabulary.

And, despite what people say about their selection process, it is impossible to overstate the impact of marketing. At least 6-10 years ago, it was incredibly common for "Bose" to be equated to good sound. In the minds of a shockingly large and diverse swath of the general public in the US, "Bose" was more or less the definition of good sound.

Measurements, at least for most people, are not a thing.

So, if someone actually gets as far as doing an in-person listening test, hyped bass, hyped treble, and simple loudness will win more often than perfect measurements. I mean, this is so well known and so common we have a name for it - "Showroom sound".

And, of course, aesthetics are a gating factor on basically any purchase, from $20 to $2M.

Basically, people really do want the best sound they can get within their budget and use case. OP's question is whether consumers actually prefer Harman sound. Their research shows they do. However, in practice, in the market, there are so many other factors at play that purely good sound (by the Harman standard or any other) is not nearly enough to achieve high sales numbers.
Best answer here.

I used to work at a regional electronic chain when I was in high school and I can confirm everything you said here.

The other thing I want to add is that, when people audition for speakers, they are at a showroom, where the salesman cranks up the SPL, people are coming in and out and you aren't listening long enough where your ears are bleeding from the accentuated high frequency (aka showroom sound). And then there really isn't a good way to do a quick AB with another speaker and since our auditory memory is short, it's hard to compare good vs bad or bad vs good.
 
Other way ´round: If they are reliable being rated best under blind and matched conditions, I would expect them to sell at least on a decent, if not dominant level, under sighted conditions in existing shops.

You are basically saying that the vast majority of consumers are choosing expensive speakers solely by the looks, which they knowingly perceive as inferior in terms of sound quality. I see no evidence for that. And I did not notice any B&W, Focal or Wilson looking so irresistible in the eyes of potential buyers, that this reverses inferior results in a listening test. If things would be that easy, Harman would have hired a few products designers from B&W and Focal, and made it to at least the market dominance they have with portable speakers.
:facepalm:
 
Who cares? I'm just glad Harmon did the research.
I'm exceedingly interested to see what the next round of speaker lines look like, particularly their higher-end stuff.
At the end of the day, this statement sums it all up.

There will always be fools in this hobby, and there will be fools with a lot of money. Ignorance is bliss, they them be blissful. There will also be con artists, snake oil salesmen and voodoo witch doctors out there, let's hope we here don't fall prey to the plague.
 
Rather the quite dogmatic way the results are treated by some as ´target curves that no-one must deviate from´ and how it is described as ´the one and only science´ in terms of sound quality.
This is part of my concern, but the doctors have helped to create situation where uncertainty and poor decades old outdated studies has become scientific fact of inaudibility or insignificance, and this is carved in stone on ASR without much...any interest to test and refresh opinions.

For example, chapter 4.8.2 "Phase Shift at Low Frequencies: A Special Case" in 3rd edition, Dr. Toole speculates how many high-pass filters (and mic) can be in production chain, adds possible room effect, and refers with some criticism few 4 decades old poor and uncontrolled studies.
Production chain does not necessarily have many...any high-pass filters so why should one design a multi-way speaker or subwoofer XO with loooong excess GD, based on combination of speculation and few poor studies? Not with reliable data and/or own tests. I know manufactures who have repeated that mistake several years.
Errors in transient dynamics due to bad timing are still ignored quite widely because authorities have said so, and common mid-range GD studies don't reveal that, though it's simple to create continuous wide range transient signal with LF fundamental, and emulate different excess group delays with FIR convolution. It's not "preference test" with Tracy Chapman etc. and single variable (spectral balance/tonality), though phase distortion can transform also to clear tonality change with transient signal.

Fig. 15. A correlation circle showing a map of the 40 most frequency used adjectives projected in Factor 1 and 2 dimensions after PCA.
1774904290215.png

These might be the variables we should listen and evaluate from the system to be worthy and respected members here. I probably continue to ignore Harman's narrowed psychoacoustics and rely on objective measurement data. Distortion is distortion no matter how many doctors try to ignore it.
 
So, if someone actually gets as far as doing an in-person listening test, hyped bass, hyped treble, and simple loudness will win more often than perfect measurements. I mean, this is so well known and so common we have a name for it - "Showroom sound".
Just like TV's, which when displayed in a store, are all set to "Vivid" mode.
 
Just like TV's, which when displayed in a store, are all set to "Vivid" mode.
It's snakey snake in the world of HiFi (and electronic entertainment in general).

For those who didn't know about these voodoo witch doctors, con artists and snake oil salesmen before you discovered ASR, you have Amir and this community to thank.
 
Also listening single "center speaker" in fixed setup is limited if the final goal is stereo listening.
In Harman blind listening tests stereo listening was also used. At this moment I can't pinpoint where I read that information, but I read it at several different places at different times, and I am pretty sure it was said by Floyd Toole himself (or was it Sean Olive?!).
 
Just like TV's, which when displayed in a store, are all set to "Vivid" mode.
That's funny. I buy all my electronics on their specs, including tv sets. The only exception is speakers.
We used to have a chain called Sound Advice that carried middle to highish-end audio gear. I spent hours searching for speakers. The B&W 800 series were in vogue, and while they looked gorgeous with the cute little tweeter tube on top, I found that they sounded surprisingly ordinary, even when powered with a large Macintosh amp. Then I gave a listen to a pair of NHT 2.5i. They were small for towers, but they sounded so, for lack of a better word, complete. I bought a pair and loved them at home as much as I did in the showroom and went back that week and picked up the matching center channel. I then upgraded to the higher models bit by bit and ended up with the 3.3s I still use today.
I almost always hate shopping...
 
situation where uncertainty and poor decades old outdated studies has become scientific fact of inaudibility or insignificance, and this is carved in stone on ASR without much...any interest to test and refresh opinions.

This ´carved in stone´ is a very good point. What is oftentimes referred to as ´scientific´ on this board, setting extremely high thresholds how controlled, blind listening tests should look like, is in my understanding a measure to prevent any other entity than Harman, to ever produce or publish any test result that would be accepted as ´scientific´, due to high costs of blind test facilities and staff to run it. While on the other hand disparaging all other tests conducted ever since as ´non-scientific´, ´biased´.

Interestingly, Harman apparently lost interest in conducting controlled tests on loudspeaker sound quality, according to Dr. Olive somewhen around 2012 (if I recall it correctly). Add these two things up, and you basically have some decades-old ´scientific findings´, based on existing products and measurement techniques of the days, which are frozen in time ever since, defended by everlasting exegesis, which must not and cannot be contested, will never be put to the test including practical findings and product solutions which evolved after the last wave of ´scientifically righteous´ products developed in the late 2000s. I personally very much would like to see some cardioids and line sources being put to a blind comparison test, not to speak of different speaker properties (you mentioned GD) in isolated testing.

The interesting question, having to do with the title of this thread, is: why? And where are these dominant products evolving from ´the science´, that was stopped some 15 years ago?

Errors in transient dynamics due to bad timing are still ignored quite widely because authorities have said so, and common mid-range GD studies don't reveal that, though it's simple to create continuous wide range transient signal with LF fundamental, and emulate different excess group delays with FIR convolution.

I agree, although I am pretty cautious when it comes down to predicting an outcome or clear audibility thresholds.

Interestingly, K+H (today: Neumann) conducted such tests, and invited recording engineers to take part, when launching their first FIR-controlled speaker in the early 2000s. If I recall it correctly, that product was named O500C, and included instantaneous implementation of GD for listening tests. The ´fullrange linear-phase mode´ was particularly interesting.

In Harman blind listening tests stereo listening was also used.

Can confirm that, they had the ability to to do that. For loudspeaker evaluation, particularly for judging tonality and comparing overall sound quality to competitors, according to Dr. Toole, mono testing was dominant, for several reasons such as better discrimination.
 
Production chain does not necessarily have many...any high-pass filters so why should one design a multi-way speaker or subwoofer XO with loooong excess GD, based on combination of speculation and few poor studies?

Not that one should necessarily be an excuse for the other, but production chain depending on the conditions can have many different HP filters in place. When I have to record and mix live concerts which incl. large orchestra, rhythm section, vocals - one does what one must in order to get a cohesive end product, and usually this involves more than a few HP filters (among others).

Other times, sure, there can be none - it's not something the consumer can know for sure.

These days we can of course we have an easy choice of types of filtering all with the click of the mouse. One day, when I have spare time (cough), I'll explore otential audible differences.. I'm much more worried these days about post-processing by media platforms. There is one concert I did which for some reason on youtube has the option for "vocal presence" which is ON unless you turn it off. I did not know of this option and it nearly gave me a heart attack, assuming this would be what people would be hearing at all times.

Oh, and as always, your input is much appreciated.
 
Back
Top Bottom