• Welcome to ASR. There are many reviews of audio hardware and expert members to help answer your questions. Click here to have your audio equipment measured for free!

"Room compensation for loudspeaker reproduction using a supporting source"

TThe section entitled "Recreating the immersive experience" in the John Beerends paper is especially relevant. For anyone who hasn't read it yet, the authors have developed a dedicated supplemental speaker with a radiation pattern of 300 degrees in the horizontal plane, a pair of which are placed behind the listener and aimed away from the listening area; i.e. towards the side and rear walls. The outputs of the supplemental speakers are delayed by 10-20 milliseconds, and their volume level is user-adjustable. I didn't notice any mention of signal processing beyond the delay and volume adjustability.
I had not heard of the paper either, it probably doesn't get a lot of discussion here because of the following highlighted quote which goes against the standard ARS dogma.

The reproduction of music is seldom improved by adding more playback channels beyond the typical two.
 
I had not heard of the paper either, it probably doesn't get a lot of discussion here because of the following highlighted quote which goes against the standard ARS dogma.

The reproduction of music is seldom improved by adding more playback channels beyond the typical two.

I don't agree with that statement either, but imo that doesn't invalidate the entire paper. What did you think of their investigation into recreating the sensation of immersion?

Within the context of wavefield synthesis at least, there is evidence that two channels results in the least amount of coloration. Francis Rumsey cites a finding to that effect in his lecture, here cued up to that part:

 
Last edited:
I don't agree with that statement either, but imo it doesn't invalidate the entire paper. What did you think of their investigation into recreating the sensation of immersion?

Within the context of wavefield synthesis at least, there is evidence that two channels results in the least amount of coloration. Francis Rumsey cites a finding to that effect in his lecture, here cued up to that part:

Thanks for the video link, very interesting.

I am a bit of a luddite when it comes to MC. I have tried it and gave up after multiple premature AVR failures drove me to vintage / DIY. Recently I have been playing around with multiple subs and manual DSP corrections and DIRAC DLBC and DBA and other DSP schemes. Currently I have switched to powerful stereo subs tightly colocated with the mains crossed with steep linear phase crossovers and listening mid-field. In addition to a relatively easy set up it allows for stereo bass, high SPL, low distortion, low extension, good time domain behavior, and a relatively coherent wave front. To me this is the best measuring and sounding system I have ever had, especially when it comes to "slam" and "dynamic contrast".

I have never been strongly attracted to the envelopment MC offers as when I go to a concert I see and hear the music in front of me so a stereo image is OK with me. I can't tell for sure but with my current system I do feel like I get some of the AE effects which have been mentioned on these boards on some songs which is different from what I have experienced before.

I am interested in this immersion system because it looks relatively easy and since most of my music was originally created as stereo it avoids up mixing schemes which I am suspicious of.
 
When I experimented with difference signals (l-r and/or r-l), I found the net result to vary quite a bit from one recording to the next. With some recordings the contribution from the difference signal speakers was way too much, and with some it was imperceptible unless I turned their volume up. So I ended up adjusting the level of the difference signal speakers way too often. Did you find anything like that to be the case in car audiophile installations? Maybe it's all good in car audio, whereas I was nit-picking it in a home audio setting.
Once I got it dialed in during the first couple weeks of setting up the system, I've not adjusted the rear fill volume since.

This is meant as a stereo enhancer rather than a surround sound upmixer, so you don't want to particularly be conscious of what comes out the rear speakers. It is much better to under-do the effect than overdo it and risk distraction. It's overdone if left- and right-panned sounds are pulled in closer to the listener than center-panned sounds. The front left and right images should stay as far away as before, but seem to be in a bigger space, and if you're lucky, wider apart from the center.

If your rear speakers seem too prominent, it's not because of the differential signal. They're probably just too loud. If you have something entirely panned the left side, then L-R equals L anyway! Again, you want rears at least 10db lower than the mains because you're simulating reflections in a larger space that will be attenuated by twice that larger path length. Intunitively, a longer delay means more attenuation, but you can play with all these parameters.

You can never get rid of all the reflections in a car, especially the side windows for pillar-mounted speakers, but you would do your best to aim them and absorb what you can of such cramped first reflections as well. The less influential your real space, the easier it is to superimpose the virtual space. Applies to the home too, of course.

Why don't I adjust the rear volume? These recordings that have more in the difference channels are simply the recordings that have more stereo width, that most need the enhancement. I configure the system to have that pleasing, non-distracting enhancement on electronic music with full panning and lots of reverberation. For a recording with less stereo width, the soundstage is expanded proportionately less. I like that consistiency. I also feel exaggerating the stereo with proportionately more is less beneficial, maybe even inappropriate, on these more center-focused tracks. I know that may be an odd thing to say as someone doing so much digital processing and generating more audio channels that aren't in the original recording. Since the listening environment is compromised, we use other compromises to make up for it.

Besides that, I don't want to think about adjusting it while I drive, I just want to listen to any type of music without feeling that it sounds wrong. Of course we all think about our sound systems, but I should keep that to a minimum while using it and especially on the road. ;)

If you search for the Gerzon's trifield thread on this site, you'll see that algorithm also works by converting stereo into mid/centere and side difference signals. It works together perfectly with differential rear fill, the content in the front and rear side speakers being similar. You can also read the white papers to see exactly how the center content is differentiated.

I do think differential rear fill or the more advanced techniques linked in OP could have benefits in the home, especially if you have a small listening room but good absorbtion materials. In other scenarious, it could be quite superfluous. I'm currently not using it at home.
 
I have a lot of problems or at least concerns with the way that this article was written, also this was not actually about the same topic, but I wanted to point out that I recognized the omnidirectional loudspeakers that was shown in the last diagram, almost certainly the Davone Mojo.
Problems and concerns in what way? It is not exact same, but I thought still relevant and others may find it interesting.
 
Problems and concerns in what way? It is not exact same, but I thought still relevant and others may find it interesting.
Yes, thank you for sharing. I hadn't seen that article before. It was interesting. Here are some initial problems and concerns:

"The swell of the orchestra reaches a crescendo, all of the instruments together creating a swirling field of sound that fills the concert hall and surrounds the listener. Anyone who has ever attended a classical music concert has probably encountered that joyous feeling of being completely immersed in sound. But most of us don’t have an orchestra at home, and a large orchestra probably would not fit in there, anyway. "

From the start, this article sets up a relative standard that many listeners may not have experienced in a genre that is likely less popular overall, and we already know that classical music concertgoers don't seem to fall into a single preference category (https://www.audiosciencereview.com/...cert-hall-acoustics-links-and-excerpts.51487/)

"Engineers have been seeking ways to re-create the immersive experience of a live music performance ever since 1877, when Thomas Edison made the first crude recording of himself reciting “Mary Had a Little Lamb.

I'm not sure that a monophonic (presumable close-mic'd) recording of speech necessarily inspired these engineers to re-create an orchestral performance of "Mary Had a Little Lamb" as an immersive experience.

"The ultimate goal has been high-fidelity audio, or hi-fi: the reproduction of sound without audible noise and distortion, based on a flat frequency response within the human hearing range."

How does one define noise and distortion (including non-linear types like IMD and possibly hysteresis, as well as temporal ones like the relatively amplitude and timing with respect to early and late reflections, critical distance, so-called "reverberation" in small rooms), also where is the flat frequency response defined, like loudspeaker near-response, surely not the listening position?

"Even moderately priced consumer equipment can process sound accurately; given that humans only have two ears, a simple stereo setup with two speakers would seem sufficient for the job."

Probably a single speaker would have been sufficient to reproduce Thomas Edison reciting "Mary Had a Little Lamb."

The authors seem to vacillate in their intended meaning of what "immersive" or "immersion" means. See the first paragraph above where they seem to suggest being surrounded by music, as opposed to the bottom of the page where they write "A sense of immersion is crucial for a satisfying musical experience."

I have to ask you whether you truly believe the latter statement is really true. I think they're conflating subjectively experiential terms like sensorial spaciousness/envelopment/immersion with higher-order processing resulting in engagement/engrossment/?another term. As an example, VR creating the sensory experience of a different environment vs art/literature/music creating an experience transcending the environment in which you are physically, where the boundaries of your physical environment "fall away" and you are, for lack of a better term, enraptured.

I honestly hope that you may have had an experience at some point in your life, regardless of what others may describe as "immersion," like I am about to describe. I was a relatively talented but mediocre musician (piano and violin) growing up. As a freshman in college, my roommate happened to have his Sony Dream Machine mono clock/radio on to Karl Haas's Adventures in Good Music program, on this day featuring a memorial to the recently deceased Nathan Milstein and at this particular time playing his performance of Bach's Chaconne in D minor. I had played the other parts of the Partita but not this movement, nor had I ever looked at or heard this movement before. Nonetheless, I was utterly transfixed, and the walls of the small bedroom completely fell away in my perceptual awareness. Everything Olive and Toole write about listeners "hearing through rooms." sound quality, comb filtering--none of that was relevant.

The next one, for me, was the BSO playing the Nimrod variation by Elgar (from a good seat). Since then, my own lens with which to subjectively view all reported preference studies has been between a Sony clock/radio and the BSO in Boston Symphony Hall.

The flip side is listening to recordings like those by The Killers, less so Coldplay or Portishead, where I can have a satisfying musical experience, even in the car or a pub (or possibly more so), without "a sense of immersion." The dynamic compression actually suits The Killers for reproduction in a car or gym, whereas Coldplay and Portishead actually seem to take advantage of dynamic compression in tracks like "Fix You" or "Roads" where they successively add and subtract layers."

Probably it would be tedious for us both for me to keep going, but please let me know if otherwise.
 
Last edited:
We already have an authoritative thread on envelopment. Getting into the same topics here risks revisiting all the same back and forth, which personally I think would be indeed tedious
 
We already have an authoritative thread on envelopment. Getting into the same topics here risks revisiting all the same back and forth, which personally I think would be indeed tedious

As someone who uses supporting speakers to enable the experience of "envelopment" in two-channel playback, imo discussions of "envelopment" are very much on topic.
 
Although bass envelopment and surround envelopment involve different auditory mechanisms, both appear to be enhanced by reducing signal correlation, introducing controlled delay, and increasing non-localizable support energy.
 
Although bass envelopment and surround envelopment involve different auditory mechanisms, both appear to be enhanced by reducing signal correlation, introducing controlled delay, and increasing non-localizable support energy.

Can you elaborate on how "introducing controlled delay" enhances bass envelopment?
 
Can you elaborate on how "introducing controlled delay" enhances bass envelopment?
My interpretation of Griesinger's paper may be overly simplistic, but he argues that low-frequency envelopment is driven by fluctuations in interaural time delay (ITD), and that lateral reflected energy causes those ITDs to shift.

So my thought is delayed energy leads to ITD fluctuations and increased envelopment or maybe I am overly simplifying what he is describing
 
My interpretation of Griesinger's paper may be overly simplistic, but he argues that low-frequency envelopment is driven by fluctuations in interaural time delay (ITD), and that lateral reflected energy causes those ITDs to shift.

So my thought is delayed energy leads to ITD fluctuations and increased envelopment or maybe I am overly simplifying what he is describing

Ah, thank you. I have no idea how to create those fluctuations if they're not already encoded in the recording as stereo bass. I assume that creating or enhancing these fluctuations was something Griesinger's Lexicon processor could do.

I have been attempting to lower the interaural cross-correlation in the bass region by using subwoofers positioned to the left and right of the listening area and using the phase controls in the subwoofer amps to put the two sides of the room 90 degrees out-of-phase (in "phase quadrature") with one another. I use stereo but this theoretically offers a spatial quality benefit even if the bass on the recording has been summed to mono. I think it was a reading of Griesinger that gave me this idea. It's certainly not the most sophisticated approach to bass energy de-correlation, but I think it enhances the perception of envelopment.
 
Last edited:
I am interested in this immersion system because it looks relatively easy and since most of my music was originally created as stereo it avoids up mixing schemes which I am suspicious of.

The impression I have (which does not arise from extensive experience!) is that the results with upmixing vary enough from one recording to the next that many people use different settings for different recordings. Nothing wrong with that if you don't mind doing so, but I would find it distracting.

As we've seen in this thread, there are several different approaches to using supporting speakers to enhance the sound quality and/or spatial quality of stereo playback. I can't speak for all of them, but I'm under the impression that they tend to be more "set it and forget it" than upmixing algorithms, and do not unexpectedly introduce artificial-sounding artifacts if set up correctly. I think they tend to have these attributes in common:

1. The frequency response of the supporting speakers is tailored to at least partially correct for deficiencies in the off-axis response of the main speakers. This tends to improve the sound quality.

2. The supporting speakers increase the amount of fairly late-onset reflection energy without any corresponding increase in the early-onset reflection energy. This tends to improve the perception of envelopment without a corresponding reduction in image precision.

3. This relatively late-onset reflection energy shifts the “temporal center of gravity” of the reflections to a somewhat later time, which in turn arguably tends to disrupt the “small room signature” cues of the playback room, and this can contribute to the venue spatial cues on the recording being perceptually dominant. In other words, the net result is a more favorable set of conditions for a "you are there" presentation.

4. If set up correctly, there doesn't seem to be any significant downside.
 
  • Like
Reactions: MKR
Ah, thank you. I have no idea how to create those fluctuations if they're not already encoded in the recording as stereo bass. I assume that creating or enhancing these fluctuations was something Griesinger's Lexicon processor could do.

I have been attempting to lower the interaural cross-correlation in the bass region by using subwoofers positioned to the left and right of the listening area and using the phase controls in the subwoofer amps to put the two sides of the room 90 degrees out-of-phase (in "phase quadrature) with one other. I use stereo but this theoretically offers a spatial quality benefit even if the bass on the recording has been summed to mono. I think it was a reading of Griesinger that gave me this idea. It's certainly not the most sophisticated approach to bass energy de-correlation, but I think it enhances the perception of envelopment.
In fact the way I will be setting this up is with Lexicon Logic 7 Music mode. This is only used to drive the rear channels. Front mains still driven with direct stereo signal. I can’t take credit for this idea, from Soundfield Audio
 
This discussion inspired me to try a simple experiment. I created a preset that sends a low-level (-19.5 dB) delayed (~15 ms) copy of L→LS and R→RS, band-limited 250 Hz–7 kHz.

Just to explore the delayed support-energy concepts being discussed here.

My initial impression is that it may add a bit of depth and spaciousness on some recordings but could just be expectation bias.
 
This discussion inspired me to try a simple experiment. I created a preset that sends a low-level (-19.5 dB) delayed (~15 ms) copy of L→LS and R→RS, band-limited 250 Hz–7 kHz.

Just to explore the delayed support-energy concepts being discussed here.

My initial impression is that it may add a bit of depth and spaciousness on some recordings but could just be expectation bias.

You might try increasing the volume and listening carefully for when clarity just barely begins to be degraded, and then set the level just below that threshold. For me that ended up being about -11 dB relative to the direct sound, as measured at the listening position, but a lot of the specifics are different so that number isn't intended to be a guideline.
 
Last edited:
What are you using for omnidirectional speakers?
This discussion inspired me to try a simple experiment. I created a preset that sends a low-level (-19.5 dB) delayed (~15 ms) copy of L→LS and R→RS, band-limited 250 Hz–7 kHz.

Just to explore the delayed support-energy concepts being discussed here.

My initial impression is that it may add a bit of depth and spaciousness on some recordings but could just be expectation bias.
You might try increasing the volume and listening carefully for when clarity just barely begins to be degraded, and then set the level just below that threshold. For me that ended up being about -11 dB relative to the direct sound, as measured at the listening position, but a lot of the specifics are different so that number isn't intended to be a guideline.
 
Last edited:
What are you using for omnidirectional speakers?
You don't have to use omnidirectional speakers, nor do you necessarily want to if you're doing a pillar mount next to car windows. I've been using coaixial speakers with about 50 degree projection in the car. Positioning and aiming have their effect, but not so much as the front speakers because of the lowpass. A lot of car audiophiles are satisfied doing differential rear fill with just the stock speakers in their rear doors or rear deck. While that may not be absolutley ideal, it's not considered that important. I do think doing the pillar mount was worth it, but this is at the level of installation that cutting big holes in your car to do infinite baffle woofers is also worth it.
In general, you have multiple sound sources, so you don't need each sound source to fill your room. If you're simulating a different listening area, how much do you want the characteristics of the actual listening area? If you want a virtual reflection from a virtual far rear wall, what's the second reflection from your real rear wall from omnidirectional speakers? Of course, you can try different things as much as you want. That kind of experimentation at home, without having to permanently install something, is plenty of fun.
 
Last edited:
What are you using for omnidirectional speakers?

Since you quoted me in the post where you asked that question, presumably I'm one of the people you were asking.

I don't use omnidirectionals for my supporting speakers. I use horns because I think their directivity control is beneficial, and because they cover the portion of the spectrum that I want to focus on.

What are your main speakers, if you don't mind me asking?
 
Back
Top Bottom