In addition, results in listening tests depend on program. Harman's elevator muzac playlist could be sufficient for basic tonality tests, but it's limited.
Choosing the program material for listening tests, and making sure all participants are sufficiently aware of how it was mixed/mastered, in my understanding is key to getting useful preference tests results. Mastering engineers who know a lot of different mixes, as well as recording/broadcast engineers judging their own, I would consider to be the most reliable when it comes to judging tonality and imaging.
I might not be aware of all playlists being used in the past, but particularly minimalistic recordings with close-mic´d (female) voices who without amplification ´would not fill an opera house´, as recording engineers like to put it, and added artificial reverb, are the least reliable to judge tonality. Natural reverb field is completely removed, and tonality of vocals depends vastly on microphone characteristics, placement and EQ.
Which explanations do you believe are stronger?
That was answered by
@kimmosto already, but I would consider:
- choice of program material
- participants´ detailed awareness for material under reference conditions
- channel layout/stereo triangle geometry
- room acoustics, particularly bass
being most important. Note, that I am not saying these factors are leading to reliable or consistent results, but might explain why vast quantities of buyers at dealerships come to opposite preference decisions compared to those in controlled tests.
Harman was one of the very few who analyzed which "signals" enable the biggest discrimination of tonal problems (=resonances) at listeners and on first place was pink noise
That is an exact description of the circle of confusion going on with these tests. If you choose ´signals´ solely by the ability of listeners to identify slight tonal changes done by PEQ filters, you lay the ground for a discrimination test, not a preference test. For the latter, you need a reference ´in the real world´.
If the discrimination test thing would be taken more seriously, I suggest white noise, highpass-filtered @300Hz, and all tests executed with headphones. I am sure that this is even more ´revealing´ in terms of identifying narrow-banded resonance effects or bell filters than Tracy Chapman. But it is equally worthless when it comes to judging what sounds right, what sounds as intended by the recording engineer, and what delivers both tonal balance and reverb tonality.
I sense a lot of criticism and bitterness from some towards the Harman research
I don´t think this is aiming at Dr. Toole or Harman´s attempts in general. Rather the quite dogmatic way the results are treated by some as ´target curves that no-one must deviate from´ and how it is described as ´the one and only science´ in terms of sound quality.
His research indicates that the average listener will exhibit preferences for flat frequency response and low coloration in speakers IF AND WHEN ALL OTHER THINGS ARE EQUAL.
Not saying this very basic finding is wrong, I can more or less confirm it from own experience. Problem is, that for example taste regarding bass is showing quite some variations between participants in any tests, and it is basically impossible to isolate the frequency response from other properties when comparing existing loudspeakers, as no two loudspeakers have exactly identical directivity and decay behavior.
So people in a controlled blind test might statistically opt for the more (anechoically) linear model, but you have no clue if more exalted bass or more or less constant directivity were decisive factors as well. On the other hand, to identify the latter as decisive, you have to do further listening tests isolating one of the properties in question, which is not easy to do.
I have undergone such tests in an anechoic chamber with perfectly linearized anechoic response. Tonal differences solely resulting from resonance and decay were pretty pronounced.
Remember that I said many people get speakers home and then wonder why they don't sound anything like they did at the dealer's? An educated customer, aware of total speaker measurements ... including spinorama, distortion and frequency response ... will be better prepared to predict whether the speaker under consideration will sound acceptable in his home environment, and will do so long-term.
Fully support your theory and do think, this is a common problem in the real world.
The problem is, that commonly people who are told to ´solely rely on measurements´, are oftentimes as mislead as the ones in the stores being told to buy the speaker which sounds most exciting with a handful of demo tracks. Reason being anechoic response is mostly irrelevant when it comes to judging compatibility between room, speaker and desired placement, as direct sound FR can be easily corrected by DSP. Directivity index, response within the early reflection windows, response over the complete listening windows, geometry of bass sources potentially exciting room modes - these things are far more important in my understanding, and I see little effort from the ´measurement-objectivist´ side to really educate people here.
And I believe it is not true to say that the majority of buyers ignore or despise measurements. If that were true, ASR would not have the volume of traffic that it has.
I don´t think that traffic and purchases of rather expensive gear are anyhow connected. Many people come to articles on ASR by googling for products, many might simply like the fact that expensive stuff which they don´t have, occasions hefty criticism here. Admittingly, some might buy cheaper gear and find confirmation that this is well-measuring, which is absolutely legitimate, but not a hint that the market is really moved by belief in measurements.
The Revel Performa F228Be won with three second-place votes and a full five first-place votes.
Sounds like a jury of reviewers voted. I don´t really see the connection to enormous sales figures, or consumers opting for a particular speaker.