• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

Crowdsourcing a blind test prize or bet

ahofer

Master Contributor
Forum Donor
Joined
Jun 3, 2019
Messages
6,343
Likes
11,899
Location
New York City
This topic came up originally in the very long PS audio thread, around the release of the new Windom firmware for their $6000 DAC, which measures just as badly as the previous firmware, despite the claims of “night and day differences” from Paul McGowan and Ted Smith.

I propose as a draft idea, that we denizens of ASR crowdsource a prize for the Golden Ears who can tell certain types of equipment apart:

-the two versions of PS Audio firmware
-DACs or amps with flat frequency response +/- 0.5db
-any two cables/interconnects that pass Ethan Winers null test to -90db
-power cords, conditioners and all non-invasive tweaks that aren’t acoustic panels/absorbers of some sort

Conditions:
-level-matched exactly
-blind to the reviewers with an X choice to eliminate difference bias
-all setup and switching videotaped for later auditing
-Their choice of music
-agreed upon high end system
-perhaps a third inexpensive choice to be selected by Amir
-90% success rate over at least 10 trials

I propose a cash prize to the charity of their choice, funded by us.

But we also have to fund the testing. If Amir consents, he would need to be compensated. Presumable so would someone else.

So. I will offer (technically “recommend a grant” from a charitable fund I advise) $5000 for the charitable gift and $1000 for trial set up, at least initially. At the very least, the losers should absorb the burden of our setup team's time.

Seeking rule refinements and additional backers.
 
Last edited:
Sure, I would back this.

If I remember right, Harman protocol for blind testing says that the switch has to happen within 200 milliseconds. That and other aspects of the design of the testing room will complicate setup.
 
This topic came up originally in the very long PS audio thread, around the release of the new Windom firmware for their $6000 DAC, which measures just as badly as the previous firmware, despite the claims of “night and day differences” from Paul McGowan and Ted Smith.

I propose as a draft idea, that we denizens of ASR crowdsource a prize for the Golden Ears who can tell certain types of equipment apart:

-the two versions of PS Audio firmware
-DACs or amps with flat frequency response +/- 0.5db
-any two cables/interconnects that pass Ethan Winers null test to -90db
-power cords, conditioners and all non-invasive tweaks that aren’t acoustic panels/absorbers of some sort

Conditions:
-level-matched exactly
-blind to the reviewers with an X choice to eliminate difference bias
-all setup and switching videotaped for later auditing
-Their choice of music
-agreed upon high end system
-perhaps a third inexpensive choice to be selected by Amir
-90% success rate over at least 10 trials

I propose a cash prize to the charity of their choice, funded by us.

But we also have to fund the testing. If Amir consents, he would need to be compensated. Presumable so would someone else.

So. I will offer (technically “recommend a grant” from a charitable fund I advise) $5000 for the charitable gift and $1000 for trial set up, at least initially. At the very least, the losers should absorb the burden of our setup team's time.

Seeking rule refinements and additional backers.

Maybe set up a regular booth or space at the big audio shows...?

Big neon sign with something obnoxious and provocative...

Maybe too hard to control the space...
 
Sure, I would back this.

If I remember right, Harman protocol for blind testing says that the switch has to happen within 200 milliseconds. That and other aspects of the design of the testing room will complicate setup.

You are thinking about it from the wrong perspective. The 'Golden Ears' types maintain the opposite - that they are better at telling the difference in relaxed, alternating, long listening sessions. We know, of course, that this makes it much harder.

They also object to the alleged fidelity-decreasing properties of ABX switchboxes. We, of course, are skeptical of that.

So switching manually, at leisure is something they think works better, but increases the odds that we get to pick the charity.
 
3cekd8.jpg
 
Last edited:
You are thinking about it from the wrong perspective. The 'Golden Ears' types maintain the opposite - that they are better at telling the difference in relaxed, alternating, long listening sessions. We know, of course, that this makes it much harder.

They also object to the alleged fidelity-decreasing properties of ABX switchboxes. We, of course, are skeptical of that.

So switching manually, at leisure is something they think works better, but increases the odds that we get to pick the charity.
That makes sense, but some people will accuse the test of being wrong or biased in some way.
 
That makes sense, but some people will accuse the test of being wrong or biased in some way.

That is inevitable. They've already dismissed a mountain of blind tests. But making it cater even more to Golden Ear beliefs means we are more likely to get takers.
 
I would contribute $ but you wouldn't get any serious takers. The reviewers and manufacturers that we would all love to see demonstrate their "golden ears" have everything to lose (financially, credibility) and nothing to gain. And these people know that already.
 
  • Like
Reactions: JPA
Conditions:
-level-matched exactly
-blind to the reviewers with an X choice to eliminate difference bias
-all setup and switching videotaped for later auditing
-Their choice of music
-agreed upon high end system
-perhaps a third inexpensive choice to be selected by Amir
-90% success rate over at least 10 trials

- Let the participants train (sighted) for as long as they want.
- p=90% success rate is is not the best way to spec, there's a difference between 9 out of 10 and 90 out of 100. Instead, spec an X% probability for the Type I error. This would make the results independent on the trial sample size (in the frequentist statistics, Type II errors->0 with increasing the sample size). Otherwise, if you prefer to consider Type I errors->0 with the sample size (likelihood inference) then use http://people.musc.edu/~elg26/SCT2011/SCT2011.Blume.pdf

Careful, if somebody with an audiophile reputation will pick your challenge, it means a significant flaw was identified. No sane golden ear will ever pick a challenge he won't have a chance to cheat.
 
- Let the participants train (sighted) for as long as they want.
- p=90% success rate is is not the best way to spec, there's a difference between 9 out of 10 and 90 out of 100. Instead, spec an X% probability for the Type I error. This would make the results independent on the trial sample size (in the frequentist statistics, Type II errors->0 with increasing the sample size). Otherwise, if you prefer to consider Type I errors->0 with the sample size (likelihood inference) then use http://people.musc.edu/~elg26/SCT2011/SCT2011.Blume.pdf

Careful, if somebody with an audiophile reputation will pick your challenge, it means a significant flaw was identified. No sane golden ear will ever pick a challenge he won't have a chance to cheat.

Yes, thanks. In my head I was thinking ten trials, and we wouldn't have time or patience for more. So a probability of randomly guessing and finding a difference 9/10 times is 2%, I think (two tails).

As for picking the challenge - yes, that's why I want to crowdsource the right tolerances (and nobody is stepping up yet) for frequency response, S/N, different types of distortion, impedance, capacitance, etc. I'm happy to give the money to charity. The important thing for me, and the doable thing, given the Golden Ears' reluctance to prove their claims, is to get the right hypothesis out there about measurements vs. listening and audible thresholds. That's really what I'm paying for. Well, that and the ability to say "prove it" in short hand the next time some pompous idiot lays the GE fallacy on me.

And, as I've said a million times, if we are careful about the hypothesis and somebody wins, we will have learned something useful. Win-win. I don't mean this to be something we will never lose, but something that will educate us, in a way we deem improbable, if we lose.
 
Last edited:
I’ve corresponded with “Archimago” and he is interested in helping!
 
Has PSA been formally invited yet ?
 
So how do you propose to run the two versions of the firmware at the same time? You'll need two identical D/A converters and there lies your problem...
 
G29-No, other than Amir’s somewhat heated challenge after McGowan’s “dog poo” nonsense.

Let me clarify something: I will definitely put up this charitable gift. However, here’s what I want out of the deal: a well thought-out statement/hypothesis of the measurement corridors within which we expect each type of equipment to be audibly indistinguishable. A separate set for each component: cable, DAC, etc.

I am determined to devise a bet that I would be happy to lose. One designed so that it *could* actually teach us something about the limitations of measurement and what we think is audible. I assume some knowledgeable people might try to find a way through the criteria. Good! Let’s find the holes in our knowledge.

I need help with that. I’m hoping some of the experts here will take a shot.

And, of course, we can carve out the two last firmware upgrades on the PSA Directwallet.
 
So how do you propose to run the two versions of the firmware at the same time? You'll need two identical D/A converters and there lies your problem...

Yes, we’d need two devices to switch. They provide, we verify.
 
Yes, thanks. In my head I was thinking ten trials, and we wouldn't have time or patience for more. So a probability of randomly guessing and finding a difference 9/10 times is 2%, I think (two tails).

As for picking the challenge - yes, that's why I want to crowdsource the right tolerances (and nobody is stepping up yet) for frequency response, S/N, different types of distortion, impedance, capacitance, etc. I'm happy to give the money to charity. The important thing for me, and the doable thing, given the Golden Ears' reluctance to prove their claims, is to get the right hypothesis out there about measurements vs. listening and audible thresholds. That's really what I'm paying for. Well, that and the ability to say "prove it" in short hand the next time some pompous idiot lays the GE fallacy on me.

And, as I've said a million times, if we are careful about the hypothesis and somebody wins, we will have learned something useful. Win-win. I don't mean this to be something we will never lose, but something that will educate us, in a way we deem improbable, if we lose.
 
You should be trying to keep alpha (p -type 1 error) and beta (p-type2) as low as possible. Pushing alpha lower and letting beta rise is just not objective. This is difficult to achieve if you choose P>/= to >.9 in a any hypothesis test with a limited number of trials and participants. An article in Stereophile a god while back noted a preponderance of trials with beta above 0.5. Those tests proved nothing much.

Given that a high number for n ( trials x participants) is effectively impossible in audio eqpt testing we have to keep below 0.9. This does of sourse mean that typical audio testing does not tell us as much as we might like. Science as reported often has very high beta, but unreported. A classic example is the reports on whether school class-sizes need to be controlled. When someone at last did a meta analysis using many such trials, IE with many many n. ? Keeping class sizes low was strongly supported.
 
Back
Top Bottom