I think I may understand what donjoe0 is getting at as I have had similar frustrations along the way.Note that the answer I gave you is far more powerful than what you ask. Instead of relying one expensive and error prone blind tests, we use threshold of hearing. Once a device's impairments land below that, then we know, with extreme confidence that listeners won't be able to detect them. That is the case because the analysis is conservative and assumes such things as masking does not exist.
It is not a matter of questioning the limits of human hearing, that seems to have been well established.
What I think some people are struggling to understand better is how do we justify the assertion that the differences between two products under test are indeed below this threshold? i.e., what make a set of tests sufficient?