• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

Better Presentation of Preference Score?

I asked ChatGPT Pro to give scores and plot the components of preference score for Clarity 66 speaker and compare it to Ascilab C8C. The results were quite a bit more understandable to my eyes:


View attachment 548895
The key to visual clarity is the normalization of each component as to best it could be.

What do you think?

Any representation has its own pros and cons, the radar plot included…

First task would be to define the dimensions.
You have decided to use the component of the score and the score itself:

PPR_ON = 12.69 - 2.49*NBD_ON - 2.99*NBD_PIR - 4.31*LFX + 2.32*SM_PIR

Second one would need to define a scale for each dimension you you’d like to plot
Then you can define the % of the scale that the device is reaching then convert to grades F->A+
The %->letters relationship does not need to be linear IMO but based on bins at least.
@amirm pointed that already.

One more point is the population to be displayed many speakers score in the 4-6 bracket.
We need more dynamic than that...

Dimensions:
  • the Score itself self-explanatory but not straight forward
The score is theoretically not bounded as LFX could go <1Hz.
If one only consider 1Hz as the “best” then the maximal score would be 12.69.
If one set the “best” to be 14.5Hz (as for the SW score), then the max score is 10.
Even using 10 as the top score we don’t really see anything above 7 and some devices have a negative score…
So the dynamic would be limited.

  • Bass extension
FX = log10(f6dB) one needs a scale.
As mentioned 1Hz is the default "best" but it can be set to 14.5Hz.
There no “worst” so one would need to decide what the bottom of the scale is.
e.g. 80Hz there are psychoacoustic and practical reasons for that but IMO a bit low especially with micro monitors which can’t play loud at LF but still provide decent overall presentation.
160Hz? i.e. AM/FM receiver could be a reasonable bottom of the barrel for “HIFI”
  • ON - axis smoothness
Assuming: On-axis smoothness is NBD_ON
From the original Score equation the best is 0dB that gives the 100% but no real bottom line
  • In room response “smoothness” SM_PIR (from 0 to 1 best) and PIR smoothness NBD_PIR (best 0dB)
  • PIR smoothness NBD_PIR Both are metrics derived from the same curve i.e. PIR which pertains to the in-room response they should be lumped together IMO. Even their current name is confusing
    SM_PIR is how the PIR is captured by the LINEAR regression. With constant directivity, some speakers could be SEEN as smooth but not measured as smooth, I am not sure I would call it that.
some info:
https://www.audiosciencereview.com/...urements-community-project.14929/#post-467858

You could ask the AI to retrieve all the scores and metrics values from the plots so far and and ask it to come up with something and more importantly justify its choices.
 
Last edited:
I asked ChatGPT Pro to give scores and plot the components of preference score for Clarity 66 speaker and compare it to Ascilab C8C. The results were quite a bit more understandable to my eyes:


View attachment 548895
The key to visual clarity is the normalization of each component as to best it could be.

What do you think?
Nice chart but completely useless at least for me. No idea which sources AI has used. How was the differentiation made? I don't think this is the REAL thing.
 
Nice chart but completely useless at least for me. No idea which sources AI has used. How was the differentiation made? I don't think this is the REAL thing.
The sources were my measurements. It computed the preference score correctly. So I asked it to produce a radar plot which it did. Then I thought it would be good to have letter scores which it produced. So assume the numbers are correct and comment on whether the presentation is good.

Here is a table format:

Score componentMeasured valueGradeInterpretation
On-axis narrow-band deviation, NBDON0.425 dBAVery smooth on-axis response with little localized ripple
Predicted in-room narrow-band deviation, NBD_PIR0.312 dBA+Exceptionally low short-term variation in the estimated room response
PIR smoothness, SM_PIR0.740C+Broad response trend is only moderately well described by a smooth straight-line regression
Low-frequency extension, f640.6 HzB+Good bass extension for a conventional passive speaker, but not genuinely full-range
Overall estimated preference score5.48–5.50BGood overall predicted sound quality, with room for improvement

I think it does a better job than I do quantifying things. :)
 
May I ask what prompts you used?
It was a sequence of prompts. First I asked it to produce the preference score. Gave it one of my measurements (which I include in my reviews) and told it about CEA-2034. That is all it took for it to create a python program to compute it! (not sure if it got the bug we found in the spec right or wrong). The rest is what I described above.

I tried it on two speakers but didn't expect it to be smart enough to think I wanted them compared so plotted both on the same graph. That is when I got the aha moment that this kind of comparison between two speakers can be quite useful.
 
Creating a graph like this in excel is also quite simple:
View attachment 548971
Maybe it's just me but I find the excel input much quicker and easier to read and understand than the graphic. You could add multiple more columns for other speakers without degrading readability, but the graphic would quickly become useless.
 
As a graphical comparison it works for me. However, should we be promoting such a simplification of the preference score as a way of comparing speakers? I know it’s all we have, but until we can include distortion, maximum SPL and compression into a scoring system I think we should encourage people to read beyond the preference score and learn to interpret a wider data set.
 
Very interesting to me to see 2 devices compared that way using AI.
I also think, we need more parameters for the comparison to be relevant, like include distortion, maximum SPL and compression, and maybe more.
Can AI be trained to read the spinorama results as a whole and give us something we can understand?
All those charts in the review still look like hieroglyph to me and without a summary from an expert, I have no idea witch device is "better".
 
Problem is that the sequence of the five performance criteria around the pentagon could alter the resulting area sizes.

See two different ways to place the given performance criteria (PC x) in my screenshot, resulting in slightly different sizes of the smaller, inner shapes:

1785842117434.png
 
If I may throw in my two cents...

I would suggest that, in addition to the typical score (and calibration recommendations!!!) for far-field (typically 10 to 12 feet and up)
and the score/calibration recommendations for near-field (typically 3 feet)—which Maiky also calculates from time to time—
a specific score and (automated) calibration recommendation could also be calculated for mid-field (e.g. 6 feet / 1.8 m or 6.5 feet / 2 m),
so that all conceivable use cases are truly covered.
 
I don't understand the chart.

1785842995375.png


First, why are there both circles and straight lines joining on-axis smoothness to in-room smoothness?

What does interpolating (by any route, circular, linear...) between these axes mean? is it a weighted combination of the two?

What does it mean that AciLab C6C gets an A+ for on-axis smoothness and an A+ for in-room smoothness but only a B+ for the combination of them?
 
Last edited:
What does it mean that AciLab C6B gets an A+ for on-axis smoothness and an A+ for in-room smoothness but only a B+ for the combination of them?
Do you mean the AsciLab C8C instead? But still – I don’t see any B+ score of it in the chart.
 
The sources were my measurements. It computed the preference score correctly. So I asked it to produce a radar plot which it did. Then I thought it would be good to have letter scores which it produced. So assume the numbers are correct and comment on whether the presentation is good.

Here is a table format:

Score componentMeasured valueGradeInterpretation
On-axis narrow-band deviation, NBDON0.425 dBAVery smooth on-axis response with little localized ripple
Predicted in-room narrow-band deviation, NBD_PIR0.312 dBA+Exceptionally low short-term variation in the estimated room response
PIR smoothness, SM_PIR0.740C+Broad response trend is only moderately well described by a smooth straight-line regression
Low-frequency extension, f640.6 HzB+Good bass extension for a conventional passive speaker, but not genuinely full-range
Overall estimated preference score5.48–5.50BGood overall predicted sound quality, with room for improvement

I think it does a better job than I do quantifying things. :)
Thanks. Aha, then it seems valuable.
 
I don't understand the chart.

View attachment 549106

First, why are there both circles and straight lines joining on-axis smoothness to in-room smoothness?

What does interpolating (by any route, circular, linear...) between these axes mean? is it a weighted combination of the two?

What does it mean that AciLab C6B gets an A+ for on-axis smoothness and an A+ for in-room smoothness but only a B+ for the combination of them?
The concentric circles are the empty chart, a bit like graph paper before you’ve written on it.

The points are the spot ratings, A, B etc. Then the straight lines connecting them are simply straight lines. There’s nothing to deduce from that in itself.

There’s no combination reading. It’s simply coincidence that the line passes near the label. It’s essentially five spot ratings connected by straight lines to aid visualisation.

Does that help?
 
Do you mean the AsciLab C8C instead?
Yes i did. Thank you

But still – I don’t see any B+ score of it in the chart.
In the part of the diagram i showed, the orange line intersects and is below the A circle for about half of its length.

Or, a related question, why is the orange line between one A+ point and the next not also a circular arc?
 
Last edited:
So why are they there?
I believe the idea behind those pentagon shapes is to provide a quick visual hint about

a.) where may be the strong points or the weak points of the reviewed device

b.) if the device’s performance appears to be well balanced or if it tends to show some special »character« instead
 
I don't think I like the use of the radar chart for this presentation. Points from the Wikipedia article sum up the reasons quite nicely:

Limitations​


Radar charts are primarily suited for strikingly showing outliers and commonality, or when one chart is greater in every variable than another, and primarily used for ordinal measurements where each variable corresponds to "better" in some respect and all variables are on the same scale.

Conversely, radar charts have been criticized as poorly suited for making trade-off decisions when one chart is greater than another on some variables, but less on others.

Further, it is hard to visually compare lengths of different spokes, because radial distances are hard to judge
, though concentric circles help as grid lines. Instead, one may use a simple line graph, particularly for a time series.

Radar charts can distort data to some extent, especially when areas are filled in, because the area contained becomes proportional to the square of the linear measures. For example, in a chart with 5 variables that range from 1 to 100, the area contained by the polygon bounded by 5 points when all measures are 90 is more than 10% larger than the same for a chart with all values of 82.

Radar charts can also become hard to visually compare between different samples on the chart when their values are close as their lines or areas bleed into each other, as shown in Figure 5.

Artificial structure​


Radar charts impose several structures on data which are often artificial:

  • Relatedness of neighbors – radar charts are often used when neighboring variables are unrelated, creating spurious connections.
    • The charts have an inherently cyclic structure; the first and last variables must be placed next to each other, even if they are unrelated.
  • Length – variables are often most naturally ordinal, better or worse, though the degree of difference may be artificial.
  • Area – area scales as the square of values, exaggerating the effect of large numbers. For example, 2, 2 takes up 4 times the area of 1, 1. This is a general issue with area graphs, and area is hard to judge – see "Cleveland's hierarchy".

Bolding mine on points I think are particularly salient.

There's also the simple issue that the preference score is not terribly useful on its own, and breaking it into a chart of its constituent values does little to ameliorate that problem.
 
So why are they there?

What are there's circles and straight lines for?
Well, that’s the way that type of chart works. A bit like bars, blocks, curves, are all ways to chart or connect different values on a conventional graph.

The circles, as I said, are the blank chart. That’s why it’s called a radar chart as it looks like a radar screen. The straight lines join the various values. It’s intended to make it easier at a glance, but it’s quite possible it simply doesn’t work for you. We all have different perceptions of patterns.
 
Back
Top Bottom