• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

"Room compensation for loudspeaker reproduction using a supporting source"

youngho

Addicted to Fun and Learning
Joined
Apr 21, 2019
Messages
548
Likes
1,010
I was browsing Soren Bech's recent publications (https://vbn.aau.dk/en/persons/100940/publications/) and ran across this interesting paper "Room Compensation for Loudspeaker Reproduction Using a Supporting Source" (https://pubs.aip.org/asa/jasa/article-abstract/159/4/3006/3386156/Room-compensation-for-loudspeaker-reproduction) follow-up to a previous publication "Reverberant Sound Field Equalisation for an Enhanced Stereo Playback Experience" (https://dael.euracoustics.org/confs/fa2023/data/articles/000407.pdf).

Abstract: "Room compensation aims to improve the accuracy of loudspeaker reproduction in reverberant environments. Traditional methods, however, are limited to improving only spectral(timbral) and temporal accuracy, neglecting the spatial accuracy of loudspeaker reproduction. Proposed is a method that compensates for both spectral and spatial properties of loudspeaker reproduction, by adding energy to the perceived reverberant sound field in a frequency-selective manner using a delayed secondary supporting source. This approach allows for the modification of the Direct to Reverberant Ratio (DRR) as a function of frequency, altering spatial and spectral reproduction. The proposed method is perceptually evaluated, demonstrating its ability to alter the perception of a primary loudspeaker without the listener perceiving the supporting source. The results show that the proposed method performs comparably to a well-established commercial room compensation algorithm, and has several advantages over traditional room compensation methods."

To summarize:
1. The authors set up a pair of B&W D3 speakers (model not specified, but you can find Stereophile measurements of one D3 model at https://www.stereophile.com/content/bowers-wilkins-802-d3-diamond-loudspeaker-measurements demonstrating very uneven directivity, which seems to be common for this range of B&W speakers) a in a stereo setup (52 degree subtended angle with each speaker 2.88m from and pointed towards the listening position) in a 3.86m x 6.08m room with T60 ~0.4s.
2. They also placed a pair of Genelec 8030 (can find ASR measurements of 8030C at https://www.audiosciencereview.com/...ds/genelec-8030c-studio-monitor-review.14795/ demonstrating very even directivity) behind the listening position at a distance of 1.79m and the same subtended angle of 52 degrees and orientation towards the listening position
3. Assessors totaled eight working in acoustics, six with regular listening training. They were not aware of the presence of the supporting speakers
4. The listening position was surrounded by acoustically transparent curtains
5. Four experimental/comparison conditions were: 1. standard uncompensated Stereo, 2. Inverse (Filtering Design Using a Minimal-Phase Target Function from Regularization, https://scholar.google.com/scholar?oi=bibs&cluster=16181673021520784068&btnI=1&hl=en), 3. an unspecified Commercial room compensation algorithm, and 4. the Proposed correction which fed the rear supporting Genelecs with decorrelated ("velvet noise" sparse noise sequence) music at a delay of 10 msec and signal level of -10 dB btw 70-500 Hz, -6 dB btw 500-20k Hz.
6. Music excerpts used 1 minute samples of "Limehouse Blues" from Jazz at the Pawnshop, "Thinking Out Loud" by Ed Sheeran, and "Prolog: La selva" from Orfeo Chaman, all stimuli matched to 72 dBa. "Subjects were able to listen for as long as required and could seamlessly switch between conditions with a short cross-fade time"
7. Preference-based results:
1781015834557.png

8. No subjects perceived sound sources behind them
9. "The DRR of the uncompensated sound field has been compared to a traditional method of room compensation and the proposed method. This analysis shows that the proposed method is able to control the DRR of a sound field, whereas traditional methods are not. The proposed approach, therefore, is likely to alter the spatial and distance perception of the reproduced content, giving a more consistent or smoother perception of distance to the listener as a function of frequency...In conventional compensation techniques, where modifications are made only to the primary source, care must be taken to ensure the phase properties within and between sources preserve temporal compactness and interaural phase differences. In the proposed method, the direct component of the primary source is left unaltered avoiding these problems while still being able to compensate for the reverberant sound field."
 
Last edited:
Thanks for the interesting find! Will add to my reading list (it is a long one). For those who are interested, the paper is available at ArXiv.
 
Not everyone has a reverb/decorrelation engine or extra Genelecs, but I wonder how close a simple DSP approximation could get: route a very low-level, band-passed copy of the front L/R to the surrounds, delay it, and use polarity inversion and/or all-pass filters to reduce correlation.
 
Not everyone has a reverb/decorrelation engine or extra Genelecs, but I wonder how close a simple DSP approximation could get: route a very low-level, band-passed copy of the front L/R to the surrounds, delay it, and use polarity inversion and/or all-pass filters to reduce correlation.
Sorry! I just realized that I forgot under #5 above that the target response for the supporting speakers was a 3dB drop from 20 Hz to 20 kHz.

But yes, I also wondered about extrapolating to surround setups, though probably would also want to EQ the surrounds to near flat at the listening position at two positions 17 cm apart, similar to what they did here.
 
A couple of years ago, I experimented with using extra speakers placed around the room to create ambience using digital delays and volume attenuation - see here. I eventually got tired of it because the software chain kept breaking. But when it worked, it really worked. Those little speakers really DO create extra ambience and improve spaciousness! I found that placing them on either side of the listener against the wall worked best. My room is wider than it is long, so "either side of the listener" is still 3.5m away from the listener.

When I did the experiment, it was a hunch. Glad to see that someone has published a formal paper.
 
A couple of years ago, I experimented with using extra speakers placed around the room to create ambience using digital delays and volume attenuation - see here. I eventually got tired of it because the software chain kept breaking. But when it worked, it really worked. Those little speakers really DO create extra ambience and improve spaciousness! I found that placing them on either side of the listener against the wall worked best. My room is wider than it is long, so "either side of the listener" is still 3.5m away from the listener.

When I did the experiment, it was a hunch. Glad to see that someone has published a formal paper.
They didn't seem to assess subjective perception of spaciousness or envelopment in either (immersiveness in the preceding report was defined as "how psychologically connected to the stimuli the subject feels"). This seemed to be more about "correcting" or compensating for directivity-related and thus frequency-dependent issues in the reverberant sound field of the listening environment resulting from speaker-room interactions, no?
 
THANK YOU VERY MUCH @youngho for finding these articles and letting us know about them, and for summarizing the test set-up so thoroughly. I had overlooked some of the specifics of the set-up in my reading of that paper.

I've been using "supplemental speakers" to manipulate the in-room reflection field since 2013, and this is the first time I've seen a relevant published study.

David Smith (chief designer of the landmark JBL Model 4430 among other things) also delved into the use of supplemental speakers, as described in a couple of posts he made on the DIY Audio forum. Unfortunately he passed away before I saw his posts so I was never able to converse with him on the topic. My recollection is that he was mainly looking at using supplemental speakers to correct the perceived in-room frequency response.

A couple of years ago, I experimented with using extra speakers placed around the room to create ambience using digital delays and volume attenuation - see here. I eventually got tired of it because the software chain kept breaking. But when it worked, it really worked.

Ime adequate reflection path length (my target was 10 milliseconds, same as in both papers) matters more than the reflection arrival directions. And there are very good reflection arrival directions in the front hemisphere, specifically 60 degrees to the left and right of the centerline. Anyway if you use reflection path length to get sufficient delay then you don't need digital delay, thereby theoretically simplifying the signal path (but maybe not, depending on the specifics). Anyway here is the backside of one of my speakers, the up-and-back orientation of the rear-firing horn being how I get sufficient reflection path length (for at least 10 milliseconds of delay relative to the direct sound) without needing a great deal of distance between the speakers and the front wall:


13-603.jpg
 
Last edited:
This is a great discussion and very applicable given my personal situation !

I will soon have a very similar system using Logic 7 and similar to our great and dearly missed friend and visionary Siegfried Linkwitz. As usual a man ahead of his time. And all credit for such a system recommendation goes to Soundfield (designer of my speaker system) for making me aware of this concept and helping me get it setup (Soundfield/AJ uses such a setup in their own personal system). Looking forward to getting it up and running!

https://www.linkwitzlab.com/surround_system.htm
 
From the aforementioned posts by David Smith, aka "Speaker Dave", posted on DIY Audio back in 2015 (I didn't find them until late 2024, after he had passed away):

"I have created a test system with fairly directional forward firing elements and separate rear firing elements. Using this in the near field I am able to adjust independently the direct and reflected sound balances. I fully expected that there would be a difference between, say, having a sensible room curve with normal directivity and dipping out the treble of the direct sound while replacing it with later arriving rear sound.

"Well, after several months of trying I can't really say that I am detecting much difference between room curves arrived at with different balances between direct and reflecting spectra. Certainly a hole in the direct response can be well filled by the right amount of later arriving energy. When taken to extremes then the spatial impression is different but the steady state curve seems to define frequency balance, at least in rooms of domestic size.

"What does this mean?

"First and foremost, we can't get too pedantic about speaker directivity and polar patterns. If an on-axis dip can be corrected with an off-axis peak, then polar smoothness is not essential. Also, energy spectrum of particular reflections don't seem to matter, as long as total response is correct, again implying that polar performance is a loose descriptor.

"Constant directivity (at least above a certain frequency) can work but it isn't the only answer.

"Now the direct to reflecting energy ratio is quite audible, especially if it can be directly adjusted. I found that equal energy between direct and total reflected energy is about as far as I would want to go on classical music and too far for pop (the spaciousness is fun but you end up preferring the precision that goes out the window).

"Flat anechoic response with falling power response is a safe recipe for a good sounding speaker. My tests (and Lipshitz and Vanderkooy's) show that the total power response needs to roll downhill and that dips in power are generally benign if the axial response is flat.

"At the same time the implication is there that a non flat axial response, allied with the right power response, can also give a well balanced sound."

And in a subsequent post he goes into the specifics of what can and cannot be accomplished using supplemental speakers to correct the in-room response :

"Most of the tests I did either shelved or dipped out sections of the direct response. Then the later arriving (reflected) response could easily be used to fill back in the steady state response. It goes the other way to: you can have a mild peak or rise in one element and correct it with less energy with the other.

"One thing to think of that wasn't immediately obvious to me. If one element has a bump, it can be no greater than 3dB to be correctable. We are not talking about serial connected filters but parallel paths. Assuming equal levels of direct and reflected sound, a 3 dB peak in one needs an infinite hole in the other to just compensate. A 4 or 5 dB peak can not be compensated.

"Back to rising or falling responses, we know that flat axial response and falling power response can sound good. Such speakers did best in the early Toole study. This does not preclude a gently rising response from being counterbalanced by a more strongly falling power response or a falling axial response being balanced by a flatter off axis power. Of course, when you are depending on off axis power to flatten your perceived response balance you need to get the room acoustics just right, but it can be done."
 
Good stuff, thanks for the Cliff notes version Duke!
 
They didn't seem to assess subjective perception of spaciousness or envelopment in either (immersiveness in the preceding report was defined as "how psychologically connected to the stimuli the subject feels"). This seemed to be more about "correcting" or compensating for directivity-related and thus frequency-dependent issues in the reverberant sound field of the listening environment resulting from speaker-room interactions, no?

I'm not sure what methods the authors used, I don't have access to that paper unfortunately.
 
I'm not sure what methods the authors used, I don't have access to that paper unfortunately.
Try @NTK's link:


In the upper right-hand corner, under "Access Paper:", click on "View PDF".
 
Something like this, called differential rear fill, is very common in car audiophile installations. The typical formula is rear left = l-r, rear right = r-l, typically high pass around 150hz, low pass around 8000hz. The you turn the rears down until they are barely heard, about 10db. Each speaker is delayed to listening position based on measurements, but then roughly 10 to 30 ms is added to taste for the rears.

Not particularly demanding on the speakers, in fact, many put the utmost into the front stage but leave the rear speakers stock.

I would say it helps a lot in small cars to help them sound more spacious, and in big vehicles it is not as needed but can be fun.

I put a system in a big old fashioned sedan with a center speaker where the radio had been, running Gerzon's trifield plus differential rear fill. Once you have your delays correct and your hi/mid aimed, that's the next step up in imaging for a car system. I will tell you, it's not a sophisticated algorithm, but that gives it consistency so it will not break your immersion like pro logic does. When you get the delays and filter parameters right it makes a subtle effect, yet the stage extends further beyond the car doors.

You can also try this home. Whether it's beneficial or how to set it up will vary by room.
 
Very interesting article, though as is so often the case it appears to be a stereo-centric method/add-on, I wonder how compaitble is would be with multichannel installations. But please in any case, when you post about an article, mention its publication year (or use a common in-text citation style like 'Brooks-Park et al (2026)', for this one). It's nice to know right away where an article sits in history. and you can make text clickable, e.g. Brooks-Park et al. (2026) so you don't have to paste the long URLs.
 
From the aforementioned posts by David Smith [...]
Was the nature of the test signal(s) described at any point? Based on my own tests and my understanding of relevant psychoacoustic research, the presence or absence of perceptually significant transient onsets (auditory events) changes the results.

Briefly: My tests were done with supplementary loudspeakers which could be smoothly faded in/out. Total power was either fully or partially compensated. I found that full power compensation sounded correct for pink noise, but music required partial compensation (or, ideally, dynamically varying compensation).
 
Was the nature of the test signal(s) described at any point?

For David Smith's informal study, not that I know of. The context was a thread on DIY Audio, so I'm assuming that he listened to music at some point.

Briefly: My tests were done with supplementary loudspeakers which could be smoothly faded in/out. Total power was either fully or partially compensated. I found that full power compensation sounded correct for pink noise, but music required partial compensation (or, ideally, dynamically varying compensation).

If you're comfortable describing in more detail, I'd be quite interested.

Something like this, called differential rear fill, is very common in car audiophile installations. The typical formula is rear left = l-r, rear right = r-l, typically high pass around 150hz, low pass around 8000hz.

When I experimented with difference signals (l-r and/or r-l), I found the net result to vary quite a bit from one recording to the next. With some recordings the contribution from the difference signal speakers was way too much, and with some it was imperceptible unless I turned their volume up. So I ended up adjusting the level of the difference signal speakers way too often. Did you find anything like that to be the case in car audiophile installations? Maybe it's all good in car audio, whereas I was nit-picking it in a home audio setting.
 
For David Smith's informal study, not that I know of.
On a second look, it's certainly implied that music was included.

If you're comfortable describing in more detail, I'd be quite interested.
Which part? The context is an upmixing algorithm I've been tinkering with for several years. I was trying to figure out how best to minimize perceived timbral change resulting from frequency shaping (shelving and treble rolloff) applied to the surround outputs. Leaving the fronts alone resulted in obvious darkening, while full compensation for total power sounded too bright with music (but correct for pink noise). Eventually, I linked the power compensation to the event detection, which seems to work pretty well.

Increased sensitivity to spectral balance of the direct sound when significant transients are present follows from the emphasis placed on onsets due to cochlear compression and other mechanisms. Sounds arriving more than a few milliseconds after a transient are quite heavily suppressed even before higher neural processing.
 
I was browsing Soren Bech's recent publications (https://vbn.aau.dk/en/persons/100940/publications/) and ran across this interesting paper "Room Compensation for Loudspeaker Reproduction Using a Supporting Source" (https://pubs.aip.org/asa/jasa/article-abstract/159/4/3006/3386156/Room-compensation-for-loudspeaker-reproduction) follow-up to a previous publication "Reverberant Sound Field Equalisation for an Enhanced Stereo Playback Experience" (https://dael.euracoustics.org/confs/fa2023/data/articles/000407.pdf).

Abstract: "Room compensation aims to improve the accuracy of loudspeaker reproduction in reverberant environments. Traditional methods, however, are limited to improving only spectral(timbral) and temporal accuracy, neglecting the spatial accuracy of loudspeaker reproduction. Proposed is a method that compensates for both spectral and spatial properties of loudspeaker reproduction, by adding energy to the perceived reverberant sound field in a frequency-selective manner using a delayed secondary supporting source. This approach allows for the modification of the Direct to Reverberant Ratio (DRR) as a function of frequency, altering spatial and spectral reproduction. The proposed method is perceptually evaluated, demonstrating its ability to alter the perception of a primary loudspeaker without the listener perceiving the supporting source. The results show that the proposed method performs comparably to a well-established commercial room compensation algorithm, and has several advantages over traditional room compensation methods."

To summarize:
1. The authors set up a pair of B&W D3 speakers (model not specified, but you can find Stereophile measurements of one D3 model at https://www.stereophile.com/content/bowers-wilkins-802-d3-diamond-loudspeaker-measurements demonstrating very uneven directivity, which seems to be common for this range of B&W speakers) a in a stereo setup (52 degree subtended angle with each speaker 2.88m from and pointed towards the listening position) in a 3.86m x 6.08m room with T60 ~0.4s.
2. They also placed a pair of Genelec 8030 (can find ASR measurements of 8030C at https://www.audiosciencereview.com/...ds/genelec-8030c-studio-monitor-review.14795/ demonstrating very even directivity) behind the listening position at a distance of 1.79m and the same subtended angle of 52 degrees and orientation towards the listening position
3. Assessors totaled eight working in acoustics, six with regular listening training. They were not aware of the presence of the supporting speakers
4. The listening position was surrounded by acoustically transparent curtains
5. Four experimental/comparison conditions were: 1. standard uncompensated Stereo, 2. Inverse (Filtering Design Using a Minimal-Phase Target Function from Regularization, https://scholar.google.com/scholar?oi=bibs&cluster=16181673021520784068&btnI=1&hl=en), 3. an unspecified Commercial room compensation algorithm, and 4. the Proposed correction which fed the rear supporting Genelecs with decorrelated ("velvet noise" sparse noise sequence) music at a delay of 10 msec and signal level of -10 dB btw 70-500 Hz, -6 dB btw 500-20k Hz.
6. Music excerpts used 1 minute samples of "Limehouse Blues" from Jazz at the Pawnshop, "Thinking Out Loud" by Ed Sheeran, and "Prolog: La selva" from Orfeo Chaman, all stimuli matched to 72 dBa. "Subjects were able to listen for as long as required and could seamlessly switch between conditions with a short cross-fade time"
7. Preference-based results:
View attachment 537927
8. No subjects perceived sound sources behind them
9. "The DRR of the uncompensated sound field has been compared to a traditional method of room compensation and the proposed method. This analysis shows that the proposed method is able to control the DRR of a sound field, whereas traditional methods are not. The proposed approach, therefore, is likely to alter the spatial and distance perception of the reproduced content, giving a more consistent or smoother perception of distance to the listener as a function of frequency...In conventional compensation techniques, where modifications are made only to the primary source, care must be taken to ensure the phase properties within and between sources preserve temporal compactness and interaural phase differences. In the proposed method, the direct component of the primary source is left unaltered avoiding these problems while still being able to compensate for the reverberant sound field."
Very interesting concept! Thank you. Will read the paper later.
 
The context is an upmixing algorithm I've been tinkering with for several years. I was trying to figure out how best to minimize perceived timbral change resulting from frequency shaping (shelving and treble rolloff) applied to the surround outputs. Leaving the fronts alone resulted in obvious darkening, while full compensation for total power sounded too bright with music (but correct for pink noise). Eventually, I linked the power compensation to the event detection, which seems to work pretty well.

Thank you, that sounds like a very sophisticated algorithm! If you are free to say: Is this something that will be commercially available one day?

Increased sensitivity to spectral balance of the direct sound when significant transients are present follows from the emphasis placed on onsets due to cochlear compression and other mechanisms. Sounds arriving more than a few milliseconds after a transient are quite heavily suppressed even before higher neural processing.

I wasn't aware of that, thank you!


Thank you very much! I was unaware of that paper as well. The paper @youngho pointed out is from April 2026, so I think it's actually the more recent.

The section entitled "Recreating the immersive experience" in the John Beerends paper is especially relevant. For anyone who hasn't read it yet, the authors have developed a dedicated supplemental speaker with a radiation pattern of 300 degrees in the horizontal plane, a pair of which are placed behind the listener and aimed away from the listening area; i.e. towards the side and rear walls. The outputs of the supplemental speakers are delayed by 10-20 milliseconds, and their volume level is user-adjustable. I didn't notice any mention of signal processing beyond the delay and volume adjustability.
 
Last edited:
Back
Top Bottom