As a manufacturer, you better be sure what you design has high chance of acceptance by these customers. The designer tuning it by ear and with some music in some room is not going to get you there.
I would assume, this is what AJ is meaning by saying, first you use measurements to get drivers, general concept, crossover and alike right, have a technically linear speaker, then you do the final judgement on compromises and ´voicing´ by listening tests. Sounds like a proper way to serve both interests to me.
The notion then that you had to choose between different compromises with respect to tonality then, doesn't work. You better work them all out.
If you take off-axis response within various windows (slightly off-axis listening window, ceiling, floor, side-walls, rear hemisphere), and directivity into account, I would not agree to this statement. The need to come to a compromise with existing concepts, is rather on the decisive side. And to be frankly honest, in terms of music reproduction, I fail to understand the point in most of compromises on the market.
Anyone who has doubts on the influence of off-axis behavior, please do an A/B comparison of KEF Q11Meta and Kii 3 (or GGNTKT M3) with acoustic recordings (preferably rich on dull reverb). If tonality and imaging are indistinguishably similar, you might want to agree to Amir´s claim that only on-axis response counts in terms of tonality, and off-axis should just be ´smooth´ (which is still undefined, btw).
That dip at 2 to 4 kHz must not be there, nor the clear resonances above it.
I don´t mean to defend this particular TAD product, but looking at the response graphs in the 10-25deg window, it becomes apparent there is excess energy in this particular band (2-3.5K). Which is a clear indication that balanced tonality on axis was not the main goal, but in this window. Not an uncommon strategy of a compromise with large conventional coaxial drivers, as they tend to have a narrow cancellation or suckout dip on axis anyways, in this case rather high, around 9K-10K. A conventional coax is always a compromise in terms of even directivity.
I remember some comments AJ was making, when the first TAD mastering speaker was introduced (almost 20yrs ago). Pointing out, he is sensitive to peaks in the 3K region, so would make sure that there is no angle within a potential listening window, which shows overshoot energy in this band. Even if the compromise involves a dip on axis.
If you conclude this should not be, don´t buy the product.
@Arindal has strong conviction that measurements are not good enough to tell the story of a speaker,
To be clear here: I am not saying, that they were not good enough, or any kind of ´unmeasurable magic´ was at play. I consider myself 100% audio objectivist and certainly agree with Amir on most of questions regarding high end gear, electronics, cables or alike.
Rather the opposite, my point is that measurements describing a very complex three-dimensional soundfield (plus time) in a room, are so complex, that overly simplified psychoacoustical models of interpreting them in order to precisely predict human perception, will inevitably fail. Just looking at the frequency response at one particular point, is in my understanding such an overly simplified model.
For those who don´t believe me - please proceed to suggested KEF vs. Kii experiment.
I thought optimization by ear is all about “voicing” that is verifying that the speaker produces naturally sounding speech?
Cannot speak for any loudspeaker designer, and methods might vary, but would assume that speech is only one aspect. Basically every sound that human voices are capable of producing in a room, as well as overtone-rich instruments like brass, and the resulting reverb pattern they leave on a recording, might be taken into account when judging overall tonality.
No of us have reference to the recorded material or the sound in the control room, but for sure we know instinctively, and without reference to the original sound, if voices sound unnatural.
I agree to the latter, but want to remark that this is rather about direct sound, not necessarily helpful to judge reverb tonality. And to the former: yes, knowing how the recorded material was coming to sound as it sounds, I would regard to be vital for judging tonality. I personally have a selection of recordings, which I, with my own ears, have heard ´in the making´. That usually involves listening to the general rehearsal in the concert hall, listening to the broadcast mixdown during the concert, and comparing it with the final product after mastering. This gives me personally a pretty good reference of how the original sounded.
I have tested hundreds of speakers where I check the audibility of deviation from flat on-axis response.
Maybe that is the explanation for different takes on judging tonality. I prefer comparing to the original concert soundfield and broadcast mixdown in the control room, your judgement is based on ´minimum of tonal deviation from flat on-axis response´. I am not saying that this is illegitimate, but I hope you are aware of the circle of confusion being at play in your case, as you have set the reference a priori from technical definition solely at one point in an anechoic chamber (real or calculated).