In principle, you are right: done properly, headphones can sound indistinguishable from real speakers in a room using digital signal processing (DSP). I'm solely listening to music this way and my speakers are used for watching TV with the family only. But unfortunately, it is not simple at all as outlined below, and this also explains why binaural (aka dummy head) recordings don’t work for the vast majority of users and recordings.
Individual perception of sound
Among other cues, the human auditory system uses tonal sound characteristics (timbre) to locate sound sources. The timbre is created by the interaction of sound waves with your torso, head and auricles, mathematically described as head-related transfer function (HRTF). Consequently, timbre is highly individual as is the size and shape of our bodies and auricles. Differences between individual HRTF can be higher than 20 dB (!) easily.
If the timbre doesn’t match the individual expectation, sound sources are perceived as being at the wrong location and too near or even inside of the head, an experience that is typically connected to generic HRTF solutions.
Headphone frequency response
Individual differences in HRTF are also responsible for the fact that different users perceive the sound from the very same headphones differently. In addition, frequency responses vary greatly between different headphone models. Both need to be taken into account, too.
Head rotation
When listening to headphones, the sound doesn’t change with head rotation, which is highly unnatural. For a small group of users it is not enough to match the individual timbre. They require head-tracking to match the expectation when rotating the head.