This ´carved in stone´ is a very good point. What is oftentimes referred to as ´scientific´ on this board, setting extremely high thresholds how controlled, blind listening tests should look like, is in my understanding a measure to prevent any other entity than Harman, to ever produce or publish any test result that would be accepted as ´scientific´, due to high costs of blind test facilities and staff to run it. While on the other hand disparaging all other tests conducted ever since as ´non-scientific´, ´biased´.
Interestingly, Harman apparently lost interest in conducting controlled tests on loudspeaker sound quality, according to Dr. Olive somewhen around 2012 (if I recall it correctly). Add these two things up, and you basically have some decades-old ´scientific findings´, based on existing products and measurement techniques of the days, which are frozen in time ever since, defended by everlasting exegesis, which must not and cannot be contested, will never be put to the test including practical findings and product solutions which evolved after the last wave of ´scientifically righteous´ products developed in the late 2000s. I personally very much would like to see some cardioids and line sources being put to a blind comparison test, not to speak of different speaker properties (you mentioned GD) in isolated testing.
The interesting question, having to do with the title of this thread, is: why? And where are these dominant products evolving from ´the science´, that was stopped some 15 years ago?
I agree, although I am pretty cautious when it comes down to predicting an outcome or clear audibility thresholds.
Interestingly, K+H (today: Neumann) conducted such tests, and invited recording engineers to take part, when launching their first FIR-controlled speaker in the early 2000s. If I recall it correctly, that product was named O500C, and included instantaneous implementation of GD for listening tests. The ´fullrange linear-phase mode´ was particularly interesting.
Can confirm that, they had the ability to to do that. For loudspeaker evaluation, particularly for judging tonality and comparing overall sound quality to competitors, according to Dr. Toole, mono testing was dominant, for several reasons such as better discrimination.