That's fair. there are conflicting explanations on this subject. Should I instead reference subjects' responses in “The audibility of typical digital audio Filters in a high-fidelity playback system ” Convention Paper, Presented at the 137th AES Convention 2014, the paper you referenced in your post "High Resolution Audio: Does It Matter"? At least we can agree that for "some reason" there is an audible difference between high resolution and filtered content, although it is not entirely understood.
The conclusions in that paper are highly contested, for instance on this very forum in the MQA thread.
I'll summarize:
1. Across all 160 trials, the aggregate result is is 56.25% correct in identifying "hi-res" vs. downsampled. This is hardly statistically significant.
2. When low-passing the signal without quantizing to 16 bits, the result was statistically insignificant. So they spun this against the digital filters, instead of investigating further.
3. They used clearly non-optimal rectangular dither, which is known to produce intersample modulation distortion, which can certainly be audible. The fix is to use a proper dithering algorithm instead, such as triangular dither which is included in just about every piece of audio software worth mentioning. The fix is not to declare CD-quality PCM audio as "bad" and insist on "hi-res" audio.
4. The difference between filters is balancing on the edge of statistical significance (talking about the 16-bit samples that the participants could identify). It also points towards the opposite of their original hypothesis. They thought a sharper low-pass filter would be more audible, but it turned out to be less audible. They spin this into something about temporal smearing and ringing in a wider frequency range, which seems like post-rationalization.
5. They erroneously think that the gap between samples imposes a limit on the time-domain resolution, leading to "grainy" or "digital" sound. Obviously this is complete nonsense, as anyone familiar with digital sampling will tell you. It's just a repeat of the "stairsteps" myth.
All the paper shows is that if you use bad dithering or no dithering at all when downsampled, some people can (with great difficulty) identify some audible flaws under optimal conditions. That's not really particularly exciting or groundbreaking.
Considering the amount of effort spent on trying to prove any kind of audible benefit to "hi-res" audio, the conclusions reached are rather damning for that entire segment of the industry.
