• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

DAC listening test files - update

I will ABX them even though I know I will just be guessings. I doubt anyone can pass an ABX on those. -85dB RMS difference, -77.5 PK metric.
Probably yes. I had a positive result on these files. Please let me answer later, I am now sitting in a garden restaurant
I am sorry but I did not find a report confirming the positive result, my memory might be wrong. All I have found is below and it is 8/16. I will try again tomorrow.

Code:
foo_abx 2.1 report
foobar2000 v2.0
2024-02-28 14:16:11

File A: DAC1.wav
SHA1: 5326127cd74c39e30bfcb9af3a3d001fc99d4022
File B: DAC2.wav
SHA1: bd5db7393f6be9c4a0c4df088e1af149625ed47d

Output:
ASIO : Topping USB Audio Device
Crossfading: NO

14:16:11 : Test started.
14:16:59 : 01/01
14:17:18 : 01/02
14:17:31 : 02/03
14:17:44 : 02/04
14:17:57 : 02/05
14:18:10 : 03/06
14:18:23 : 03/07
14:18:42 : 03/08
14:19:00 : 03/09
14:19:13 : 04/10
14:19:32 : 05/11
14:19:45 : 05/12
14:19:58 : 06/13
14:20:17 : 07/14
14:20:30 : 08/15
14:20:43 : 08/16
14:20:43 : Test finished.

 ----------
Total: 8/16
p-value: 0.5982 (59.82%)

 -- signature --
60a7066e76e516a33569282f03a15c3fd2606f68
 
can we conclude this ripple shows the main difference in the recordings?
Let me also show it in log scale:
/and good night till tomorrow/

D10s-DX5_ripple.png
 
The amplitude seems to surpass 0.1 dB only north of 30 kHz. I would be very surprised if this has any audible effect. We will see how this whole comparator thing works out. Maybe it will be useful to test at some point.
Yes, there is that. But on the other hand if the passband ripple is the only thing keeping thse from nulling to better than -110dB, what else could it be? But that assumes someone can pass the ABX. I thought Pawel had done that but now it looks like he is not sure if he remembered correctly. So someone would have to pass it.
 
But that assumes someone can pass the ABX. I thought Pawel had done that but now it looks like he is not sure if he remembered correctly. So someone would have to pass it.
Yes, I think we should start with that :)
 
Yes, I think we should start with that :)
I am of no use for that. I am willing to try but I have done the training on the ABX of those files and cannot hear any difference at all.
 
Yes, there is that. But on the other hand if the passband ripple is the only thing keeping thse from nulling to better than -110dB, what else could it be? But that assumes someone can pass the ABX. I thought Pawel had done that but now it looks like he is not sure if he remembered correctly. So someone would have to pass it.
There are some subtle “clicks” in the deltawave null that are obvious on the spectragram delta, though I couldn’t hear them in an ABX myself.
 
There are some subtle “clicks” in the deltawave null that are obvious on the spectragram delta, though I couldn’t hear them in an ABX myself.
Can you show me what you mean. I don't read spectrograms very well. All I really see on the delta spectrogram is the wave of the passband ripple. If I zoom in I see some flecks of color, but I have no idea what that means. Is it something besides noise? Are you talking about something else?
1786230556052.png
 
1786241933907.png

Sorry maybe I got confused by a change of topic, this is what I see on the original pair of files pma posted.
 
@oleg87 That seems to be correct speaking about the original files file_1 vs. file_2 and it reflects filter attenuation of DAC/ADC path above 40kHz. We need to be quite accurate in commenting the files as we now have 2 different tests, as is mentioned in the updated post #1. Post #1 is always the basic post, it can be edited and should be checked, because it is easy to get lost in the long chain of other posts, as the thread is getting longer.
On the other hand, the summarized spectral view does not tell much about instantaneous differences. It is a macro view, not a micro view. Resolution in a macro view is questionable and do not forget that the plot over the whole music sample (the image) must have data points reduced to match the display pixels, the number of pixels on X axis is not infinite and is much lower than the number of samples in the music file.
 
Hi! Good idea for a thread, considering the discussions currently floating in public opinion (also referring to Jim Lill's video on preamp influence on sound, which despite somewhat imprecise methodologically speaking, raises useful points).

Had some fun and very obsessive moments ABX testing this. I mix and master music for a living for quite some years now, so I was very interested in taking the test, since critical listening and blind tests are daily tasks. I find it important to keep a constantly updated sense of where the threshold of perception sits at for a given audio technique, since science informed decision-making necessarily looks different in either side of said threshold. For example, 320 kbps mp3s are consistently identifiable in ABX testing by well trained ears (with the caveat of most differences being present in the high end, where listener age matters most), but mostly undistinguishable for most people. Therefore, knowing exactly how small the difference is for most program material, I have no problem sending clients 320 mp3s for early mix evaluation. This forum has been important for solidifying knowledge on much of what I'm describing.

On to the tests. I did one round of training on foobar to get a general impression of the difference between files. It was clearly obvious, on first listen, that the difference would either only be relevant to the upper percentile of listeners or be mostly undetectable. Starting the test, I set up a 2 or 3 bar time range, making sure playback always restarted when switching X/Y (along the time domain, program material with physical instruments has inifinitely more timbre variation than electronics). After listening through most of the spectrum for differences, including low end transients, which can sound longer or less sharp when a lot of group delay is present, as well as imaging, I decided to focus on the very high end of the spectrum. That's where I expect most differences to be. The ride cymbal is perfect for that - sharp transients with complex frequency and time-domain information + lots of high end content. I believed I was noticing an extremely small difference, where on file B the ride's transient was more solid but the cymbal's tail had a touch less high end. Ran 4 tests, spaced out over a few hours, since the test is quite exhausting. Results (foobar files listed below): 9/16, 8/16, 10/16, 11/16, 10/16. After finishing the first test I was quite tired and decided that I should take some breaks during the next few tests, since the last few trial rounds always felt less precise. I kept feeling like I was getting progressively better at the test and learning the program material, yet the results never reached sufficient statistical significance, so I eventually called it a day. No matter how much I wanted to pass a test, even if I had reached 12/16 on a single one, it wouldn't matter much without being able to consistently replicate it.

In frustration, today I decided to try a fifth test focusing on different sounds. Looked around for differences and my brain decided that I could hear the side channel (the difference between L and R channels) "compressing" ie losing amplitude during loud vocal passages in a specific section and ran a test focusing on that. The result was a glorious 3/16.

Final remarks: Yesterday I was determined to believe I could hear a difference in the ride cymbal, but that I was being held back by fatigue. In full honesty I have serious doubts I ever perceived a difference present in the files, no matter how clear that seemed at the moment. I am fairly sure that hallucinating differences is extremely common on untested critical listening of this level and, in fact, think it might unfortunately have a much bigger role in audio quality judgements on an everyday basis than actual testing with solid methodology. Most of my conclusions stemming from this test are not about DACs, but about methodologies and ethics in audio testing. Uncontrolled confounding variables simply have way too large a magnitude of influence to be overcome by listening experience in many situations eg: trained listeners trying to listen beyond room influence when comparing speakers in different spaces. And, as timbre and psychoacoustic literature has shown, within confounding variables, context specifically is show to have an unexpectedly large influence in the way we perceive timbre.

I also purposefully didn't search for very low amplitude moments in the test signal, where noise floor or low amplitude conversion artifacts could be present. If you have to wait for the last few seconds of a fade-out or reverb tail to distinguish two pieces of equipment, I don't believe that can be pointed out as a timbre difference, considering the musical program material humans use audio equipment for. Unless that low level artifact clearly sounds worse or less faithful, that's in the realm of testing lab equipment in my opinion.

I am currently away on holidays, so I listened on my mobile setup (AKG K712 Pro connected to a JCALLY JM6 Pro). Will test again in the studio with a better setup. It's possible that this DAC can't reproduce the files faithfully enough. I tried the test on my Bose Quietcomfort NR headphones (using the JCALLY and a cable), and the lack of definition in the high end, either from the ADC/DCA, DSP processing, or the headphones themselves, made it impossible to even approach the ride cymbal differences - the high end sounds much more imprecise and scattered in time. To finish off, assuming my results are valid, I still would really like to abx test signals routed through multiple ADC/DAC stages, as that's something I strive to avoid in my setup.
 

Attachments

Last edited:
@pedrojoaquimborges : Thanks for testing, for your efforts and nice post! Those multiple results seem to give a proof that you were able to hear a difference.
 
@pedrojoaquimborges : Thanks for testing, for your efforts and nice post! Those multiple results seem to give a proof that you were able to hear a difference.
Oh, do they? I am not sure how the p value calculations change when doing a much higher number of trials (i did 64 on yesterday's test, I assumed I needed 48 correct answers to reach statistical significance).
 
8/16, 8/16, 7/16, 8/16. So, nothing. My memory about having positive result two years ago was wrong, all we always need is data.
Thanks for doing that again. I will do an ABX this morning. In general, if someone gets 7, 8, 9 out of 16 the first time, is there any point in doing more ABX if it just feels like guessing?
 
Oh, do they? I am not sure how the p value calculations change when doing a much higher number of trials (i did 64 on yesterday's test, I assumed I needed 48 correct answers to reach statistical significance).
Can you please try with the second pair of files? We are looking for someone who is good at this!
 
Oh, do they? I am not sure how the p value calculations change when doing a much higher number of trials (i did 64 on yesterday's test, I assumed I needed 48 correct answers to reach statistical significance).
I think you need 45 to reach p<0.05 with 48 reaching p=0.018 for N=80. So the numbers say "significant", if you assume that 5% is a good threshold.
 
Oh, do they? I am not sure how the p value calculations change when doing a much higher number of trials (i did 64 on yesterday's test, I assumed I needed 48 correct answers to reach statistical significance).
If we get your 10/16, 10/16 and 11/16 results, then we get 31/48 which makes p = 0.02973. Provided that we omitted the 3/16 and 8/16 results. That is why I say it seems that you were able to hear a difference.
 
Okay, so I decided to go beyond the ABX testing methodology and "cheat" a little. Loaded up the files in my DAW. First of all I'm surprised at how similar the conversion is in the time domain for high frequencies. Null test clearly shows some small differences, one of them being the ride cymbal, so I decided to find a section where the ride signal is most present in the null test (and therefore there's the most difference between files). Ran the ABX test for just that one section and decided to do it in 4 segments of 4 trials with some breaks in between to try to reset my hearing. Result: 12/16.

But at this point, this feels like desperately trying to demonstrate the difference between files. I had to condition the test variables so much to chase this result that I still seriously doubt it is useful. AFAIK the converter's clock could have drifted more than usual in that moment of the test signal and all I'm hearing in the null test is a very slight timing difference.
 

Attachments

Okay, so I decided to go beyond the ABX testing methodology and "cheat" a little. Loaded up the files in my DAW. First of all I'm surprised at how similar the conversion is in the time domain for high frequencies. Null test clearly shows some small differences, one of them being the ride cymbal, so I decided to find a section where the ride signal is most present in the null test (and therefore there's the most difference between files). Ran the ABX test for just that one section and decided to do it in 4 segments of 4 trials with some breaks in between to try to reset my hearing. Result: 12/16.

But at this point, this feels like desperately trying to demonstrate the difference between files. I had to condition the test variables so much to chase this result that I still seriously doubt it is useful. AFAIK the converter's clock could have drifted more than usual in that moment of the test signal and all I'm hearing in the null test is a very slight timing difference.
Just out of curiosity: Could you name the time stamp?
 
Back
Top Bottom