• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

A Broad Discussion of Speakers with Major Audio Luminaries

Which is also known as Beranek's law.
View attachment 523516
For my last DIY speaker build I decided to "get serious" and learn VituixCad and I measured each driver and took them outside on a turntable and compiled the data and imported it into the program and started working on fine tuning the design for flat on axis response, good off axis response, and good directivity. After hours of changing one thing to make one thing better and one thing worse and none of them particularly great I gave up and bought some Neumann KH 310's. Unless you have actually tried to build a speaker that measures like Neumann or Genelec or the like I don't think you can appreciate how difficult it really is. Instead of being amazed by how expensive the speakers were I was amazed at how much engineering must have been done and the felt like they were kind of a bargain :) DIY speakers can be fun, especially for more exoctic concepts, and can sound very good but in most cases the well engineered speakers are very tough to beat.... and they sound great.
 
Most people I know who love the built quality and/or premium feeling of their speakers just pet them any time they are near them, they can't help themselves otherwise.
!! I have to pass right by my RF speaker on the way to the washroom a few times a day. I have had to consciously restrain myself from petting it with my hand while I say "Hello Baby, I'll rock you a bit later". Don't want to wear a shiny spot on the beautiful matte finish.
 
Last edited:
So there is some evidence here that blind vs sighted results are not correlated.
If true, that would mean (as I said earlier) that actual sound has no impact on how much you like a speaker when listening to it sighted. You should therefore buy your speakers on looks alone.
Wouldn't go quite that far. The first sample did show some correlation, the second not so much. So for those limited samples and folks tested, perhaps SQ was good enough to be the determinant, in other instances looks made things muddled. Currently, the market uses both sound and looks, that's unlikely to change soon ;-).
It would be interesting to know all models and measurements of the samples, as well as whether tests of more modern optimized on/off axis designs would fare any different.
 
Notice that he writes "the extent of the bias," not "the existence of the bias." The bias is always there, varying only in extent, and only in ways we cannot know or correct for.

’To the extend’ doesn't rule out that extend starts from zero, or at least, neglectable.
 
True, but if that’s what @Floyd Toole meant, it’s a glaring departure from his characteristic unambiguous prose.

Don't know how you can conclude from "the nature and importance of the non-auditory factors to the individual listener" that these factors can never be that low that they don't completely throw off judgement of sound (as claimed by: "sighted listening is completely and totally unreliable" or "can't be relied upon to be accurate in any given case"). As far as I'm aware of, Dr. Tooles work doesn't even contain specific research on the nature and importance of non-auditory factors for different types of listeners, or the physical or mental limitations of different types of listeners to handle sighted bias (and then I mean very specific tests that control for extraneous variables, no assumptions derived from his preference tests).

More practical, do you think can you design a sighted loudspeaker listening test where you built in some visual queues that don't allow me to call out some speaker has a 3dB boost in the high mids? Or when you let people listen to a 10m high PA line array, do you believe they won't be able to observe that system has no sub bass notwithstanding its size?
 
Last edited:
The limited data presented in Flyod Toole's book clearly shows a correlation between blind and sighted results. Here is one relevant figure illustrating this:

View attachment 523357
If there was zero correlation, it would mean that blind tests were useless for evaluating speakers that are to be listened to sighted in the home. It would mean that sighted bias is so overwhelming that the actual sound has no impact at all on sighted listeners' perception of the sound. It would mean that, for maximum enjoyment, speakers should be purchased on the basis of non-sonic factors such as looks and cost only. Measurements (of the sound) and blind testing would both be useless for determining how listeners reacted to the speakers when they make sound, only non-sonic factors would matter.

Of course this absurd situation is not the case. The data clearly shows that results of blind and sighted listening are correlated, as basic common sense demands. The expert opinion of Dr. Floyd Toole, in his "final word" on the subject (quoted at greater length above), is that "Sighted evaluations introduce the possibility that opinions will be biased by non-auditory factors. The extent of the bias depends on the nature and importance of the non-auditory factors to the individual listeners."
[to the general reader] I hope that readers can understand that I'm getting a bit frustrated with this misrepresentation and misinformation campaign that certain individual individuals are pursuing here.

I see two major problems with the information from Toole’s experiment that MarkS is using as evidence of “obvious correlation” between sighted listening and attributes of the sound waves:
  1. Some of the sighted listening result is actually blind listening (speaker 1 vs 2). The oft-repeated claim, that the consistent score difference between 1 and 2 blind vs sighted tells us that the sighted ranking must be based on the sound waves, is a mistaken claim, because they are blind both times (compared to one another). (Also, as an aside, note that there is no difference in their relative scores….look at the error bars.)
  2. If the sighted listening involves contextual information that happens to correlate with measurements, then of course sighted listening will correlate with measurements, but not because sighted listening is giving insight into the sound waves themselves, but in fact, through the contextual proxy. A simple example: if one speaker looks like it has more bass, eg larger, and actually does have more bass measured, then sighted listening could be correlated to the measurement because the visual information (contextual proxy) is biasing sighted listeners to expect more bass, and not necessarily because they are detecting it in the sound waves. We would firstly need to control for the proxy, before citing the result as evidence that sighted listening is detecting the difference in the sound waves.
And points 1 and 2 above were both present in the Toole data that MarkS cherry-picked as his evidence. This is very disingenuous by MarkS. How do I know it isn’t just honest oversight by MarkS? Because I have pointed dome of these errors in his analysis out to him when he has used it in the past, link. And he is very intelligent and well-versed in science, so it would have been easy for him to grasp and agree, and retract. But instead he ignores and proceeds to repeat it.

If we want to validly test if sighted listening is correlated to the sound waves themselves, then the contextual non-sonic information would need to be carefully selected so as not to create expectations that correlate with measurements.

Which leads me to the next post.
 
I am using "correlated" in the sense meant in statistical analysis.

Eyeballing the numbers in Toole's charts, I get 6.75, 6.8, 6.3, 6.25 for the blind ratings and 7.5, 7.7, 5.95, 6.45 for the sighted ratings. Plotting blind vs sighted (blue points) and the best least-squares linear fit (orange line), I get:

View attachment 523466

A standard measure of how "correlated" these two sets of numbers are (how close the blue points come to being exactly on the organe line) is the Pearson correlation coefficient r, which has a maximum value of +1. A value of r above 0.8 is typically regarded as a "very strong" correlation.

For these numbers, the result is 0.95.

That is what I mean when I say that the blind and sighted results are "correlated".
If we are to use this dataset in the way you have, it is important, don’t you think, to consider whether the correlation is actually due to a coincidence between measurements and non-sonic factors.

I have considered that possibility and I believe that, actually, there is such a correlation, as follows.

Based on non-sonic sighted factors, which speakers might be expected to score well and less well?
  • Speakers 1 and 2 (visually identical), are likely to score well, being premium priced, large floorstanders, stylishly crafted of highly polished wood, and European.
  • Speaker 3 is likely to score relatively less well, being small and cheap, made of plastic.
  • Speaker 4 is large and premium, so likely to outscore speaker 3, but from a US competitor company to Harman. Since the participants were all Harman employees, one might expect them to be biased against giving it a high score. It would be reasonable to expect it to score between 3 and 1(2), if sighted bias is influential.
  • Result: 1(2) to score highest, then 4, then 3.
Okay, let’s look at measurements. Based on measurements, which speakers might be expected to score well and less well?
  • Speakers 1 and 2 have the most overall balanced FR from 200 Hz-20 kHz, with speaker 2 showing 1-2 dB of emphasis around 2 kHz and some decline in the mostly-inaudible top octave.
  • Speaker 3 is nice and smooth above 2 kHz, and a gradual roll-off of -5 dB at 400 Hz relative to 2 kHz. This significant dip in its midrange and overall tilted-up response should reduce its blind preference compared to 1 or 2.
  • Speaker 4 measures a nice even balance from top to bottom, but with a 5-6 dB dip around 1 kHz and one octave wide. It is also relatively narrow in dispersion at all frequencies.
  • Result: 1(2) to score highest, then 3 and 4 less preferred (I hesitate to predict between 3 and 4).
Uh oh. This indicates that there is a correlation between measurements and sighted non-sonic factors. It would therefore be incorrect analysis to conclude that sighted and blind listening preferences are correlated because of the sound waves. It may very well be the proxy causing the correlation, not the ability to hear sonic attributes sighted.

MarkS to note. Matt to note. Sigberg to note. Anyone who thought Mark conducted a persuasive analysis….to note.

Let me finish with Dr Toole’s own words when discussing the results of the test that Mark has presented without context: Again, it seems that the listeners had their minds made up by what they saw and were not in a mood to change, even if the sound required it. The effect is not subtle.

"Summarizing, it is clear that knowing the identities of the loudspeakers under test can change subjective ratings.

  • They can change the ratings to correspond to presumed capabilities of the product, based on price, size, or reputation.
  • So strong is that attachment of “perceived” sound quality to the identity of the product that in sighted tests, listeners substantially ignored easily audible problems associated with loudspeaker location in the room and interactions with different programs.
"These findings mean that if one wishes to obtain candid opinions about how a loudspeaker sounds, the tests must be done blind.” (his emphasis)

This is why I think it is plainly mischievous to cherry-pick one of Toole’s graphs and draw a conclusion that falls apart when analysed (see above discussion), a conclusion which is effectively the opposite to the author’s own conclusions from his own work. And especially bad faith considering that MarkS is a deeply experienced research scientist, by his own admission.

Not to mention the fact that I have raised many of these issues with Mark previously, see above link. So…..what’s going on, do you think?

cheers
 
So, if the listener doesn't like how the loudspeaker looks could s/he ever be satisfied with the sound?
 
MarkS said:
The limited data presented in Flyod Toole's book clearly shows a correlation between blind and sighted results.
MarkS said:
So there is some evidence here that blind vs sighted results are not correlated.
This reminds me of an old joke, popular among physicists and mathematicians.
Statistics for biologists: "Take two rabbits ..."
 
So, if the listener doesn't like how the loudspeaker looks could s/he ever be satisfied with the sound?
Ideally, no looks at all (at least at the sides)

Glorious flush-mounting and one has the best of every world possible for speakers :p
 
Has anyone ever done sighted testing showing different speakers but using the same hidden speaker for the sound? Might say some more about sighted testing. If nothing else it would be interesting psychologically to see what people say after. How many would say they heard a difference even if they didnt, how many actually believed they heard a difference, and how many would say no difference.
There was a wall of speakers with a speaker-switcher box in the big demo room at the store I worked at.
You had to be careful demoing speakers for customers, they had very little idea which one was playing as we switched between each model. If you got ahead of them, the customer would start gushing about the B&W playing when instead it was an Infinity, or whatever.:eek: Which was bad for the sale since customer confusion over what they are hearing is counter-productive unless the salesperson is controlling the narrative. I admit I got confused a few times as to which speaker I selected with the switcher.
 
There was a wall of speakers with a speaker-switcher box in the big demo room at the store I worked at.
You had to be careful demoing speakers for customers, they had very little idea which one was playing as we switched between each model. If you got ahead of them, the customer would start gushing about the B&W playing when instead it was an Infinity, or whatever.:eek: Which was bad for the sale since customer confusion over what they are hearing is counter-productive unless the salesperson is controlling the narrative. I admit I got confused a few times as to which speaker I selected with the switcher.
This brings back memories of my early comparison tests in which local audio salespeople would bring loudspeakers for double-blind comparison evaluation, and participate as listeners. Some were very good listeners. However several of them had to voluntarily retire from listening duties because what they learned prevented them from selling the most profitable loudspeakers in the store, or they didn't sell brands that they could wholeheartedly support.

In the situation you describe, the "fair" solution for customers would have been to hang an acoustically transparent curtain over the wall, find out which products the customer liked the sound of, open the curtain and let him/her choose the one they like the look/price of. Not perfect by any means, but a step in the right direction.. The influence of brand alone for average consumers is enormous. Even in car audio, a three-pointed star on the steering wheel was a huge bias, when it had nothing to do with the audio system.
 
This brings back memories of my early comparison tests in which local audio salespeople would bring loudspeakers for double-blind comparison evaluation, and participate as listeners. Some were very good listeners. However several of them had to voluntarily retire from listening duties because what they learned prevented them from selling the most profitable loudspeakers in the store, or they didn't sell brands that they could wholeheartedly support.
I'm probably like most people, I have never been able to participate in a properly controlled double-blind speaker comparison. Your research makes me realize how poorly controlled all of the events I participated in were. :D And how important it has been for me to improve my ability to measure and interpret the data.
In the situation you describe, the "fair" solution for customers would have been to hang an acoustically transparent curtain over the wall, find out which products the customer liked the sound of, open the curtain and let him/her choose the one they like the look/price of. Not perfect by any means, but a step in the right direction.. The influence of brand alone for average consumers is enormous. Even in car audio, a three-pointed star on the steering wheel was a huge bias, when it had nothing to do with the audio system.
Yes! Except the goal was to upsell. We had the gorgeous pair of B&W 801 with Levinson 20.5 monoblock amps in the main listening room. A curtain would have interfered with the ability to sell a midrange B&W over other models in the room with the wall of speakers. :cool:
 
The corollary to Beranek's Law is something like: "the sound quality increases in proportion to the effort expended in achieving a fine polish".
For consumers this would correspond to: "the sound quality increases in proportion to the efforts in reading/watching internet/magazine reviews"
For prosumers (aka ASR readers) this corresponds to: "the sound quality increases in proportion to the efforts put into measurement comparison"
In reality, some diy enthusiasts design better speakers than many professionals. And some highly regarded commercial speakers are rooted in diy designs.

On another note, I would ague that the worst sighted bias for tech-savy consumers comes from printed measurements and their (miss-)interpretation.
Correlation between loudspeaker measurements and perceived sound quality may be fairly well understood for frequency-response on- and off-axis.
However, room interaction of certain dispersion patterns (wide/narrow, constant directivity, cardioid) and their effects on sound quality aspects seem less clear.

When it comes to distortion, especially HD can be highly missleading. While ASR members love low THD plots, many also love high THD speakers (e.g. D&D 8C, Kii Three). Maximum SPL with respect to a constant THD limit over frequency basically makes no sense at all. Confusion starts with the combination of all HD orders into one THD value despite the fact that higher HD orders are more audible. A constant limit also ignores the upwards masking effect of the base frequency, which changes with frequency. These aspects comprise narrow-band signals like the sine-sweeps used for testing. More masking will occurr with actual music. Nevertheless, a perceptionally correct (T)HD limit/representation for narrow-band signals would be a good step forward.
 
Back
Top Bottom