Surround Sound in Headphones
soundguys.com
soundguys.com
Anyway, fun fact: OpenAL Soft has an implementation of this. If any games you play use it (Minecraft notably does) you can get 3D audio by just setting 'hrtf = true' in your OpenAL config.
Isn't that the idea behind Ambisonics [1], which OpenAL Soft is based on?
Yes, using even third-order Ambisonics is not as precise as rendering all point sources (and non-point sources...) exactly through the HRTF. But it is an existing standard you can use today via OpenAL Soft.
How have we not done better since? How is this not just in every video game now?
https://ftp.acc.umu.se/mirror/media/Oakvalley/soamc/000/MP3/...
Got more of similar stuff?
https://www.youtube.com/watch?v=I7U5pQZJYKM
(Sound of Games: "The Last Ninja - Pure Meditation" - HQ)
How so? I always wondered, why I didn't here more of this kind of recordings, since I heard a (fantastic) demo in the early 1990's.
I was also surprised to see no mention of the Web Audio API implementations of HRTF and multi-channel ambisonic decoding down to binaural in the browser. This may well be today's most sophisticated approach to spatial headphone-driven audio.
Can't speak for the sound quality of Waves NX, as I have it turned off.
Curious, here are some related links I could find.
- MDN: Web audio spatialization basics
https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_A...
- Implementing Binaural (HRTF) Panner Node with Web Audio API (2015)
Article - https://codeandsound.wordpress.com/2015/04/08/implementing-b...
Code - https://github.com/tmwoz/hrtf-panner-js
- WebGL / Web Audio API interface for listening to the CIPIC HRTF database
https://amp.reddit.com/r/headphones/comments/aouajt/what_hea...
The alternatives are:
- record a real sound space using a binaural mic head, then simply use 2-channel output. Head tracking not supported, HRTF is physical not modeled/virtual.
- build a 3D virtual sound space. Apply the HRTF in post-production with output to 2 headphone channels. Head tracking can't be supported, but there's a lot more flexibility.
- encode a potentially large number of discrete channels and their positions in 3D space (including reverb and room dynamics) and let the listener's system apply the HRTF. Would support head tracking, but the audio format would be super complex, as would be the rendering software. Reverb calculations in particular could eat up a lot of CPU.
- render in post-production to a discrete number of channels that will be placed in known positions around the listener. This supports theaters and home surround setups, can be implemented in headphones using an HRTF, and would support head tracking - all at the same time. The only drawback is that in headphones, it sounds like speakers around you. Because that's exactly what the HRTF is doing with the input signal.
It's easy to see why the last option is what we've landed on. It's the most flexible for the widest number of scenarios, it's a straightforward standard, and it's not too hard to emulate using an HRTF.
This does seem like a lot - but at the same time, it doesn’t seem like much compared to the advances in graphics over the past decades? I feel like if audio had kept up with video, we should have an equivalent to Physically Based Rendering / raytracing by now...
I think that's the issue as well; in a room with speakers all around, the sound has space to move and bounce about and merge so the individual speakers are less obvious, but in headphones it's pretty much a straight line into your ear. Might just need a bit of fine tuning and partial merging of multiple channels / directions.
Previous Dolby standards worked as you describe, but Atmos is most similar to your alernative 2.
Atmos has up to 128 audio tracks which are each positioned and moved within a virtual 3D space.
The Atmos reciever then processes this 3D space and outputs a number of channels tailored to your specific speaker set up, wherever they are positioned (and however many* speakers you have). I don't know if Atmos for headphones currently includes HRTF processing, but there is no technical reason it couldn't.
*(I believe state of the art limit for home systems is 64, but this is a hardware issue, rather an a limit of the format)
Head related transfer function is the.. modeling of how sound waves bounce of your body (shoulders, ears, breast). And it's a part of how we determine where a sound is coming from. Other part is the timing, as in the difference of phase of sound coming in one and in the other ear. Amplitude(volume) is also a part, probably, but we are much more sensitive to difference in phase then amplitude.
If you want to start researching in these things; the start is probably "impulse response" as well as.. a lot really (systems and signals is a heavy book that i should read one day). Just remember that everything is a spring.
As for practical modeling in 3D games and such. The problem is almost the same as global illumination. You would need to model the whole environment and how it responds to sound (absorption, reflection, idk probably even the speed of sound in various materials to be accurate). A great youtube person has a couple videos on whitepapers about this [1-2](the 2 is more practical). Valve bought a company that played with these things, but idk if anything came out of it (iirc it is part of the steam framework now).
[0]https://www.youtube.com/watch?v=VW-W3A2l5UE
[1]https://www.youtube.com/watch?v=DzsZ2qMtEUE
[2]https://www.youtube.com/watch?v=Mx8viOFKiIs
PS Headphones are better then 5.1/7.1 for surround sound, mostly because of room acoustics.
On mobile it looks the gaming version of Immerse is only available for headphones made by a few manufacturers. Which version should I choose for AKG or Sennheiser?
Also I'm not entirely certain why I'd want a subscription as opposed to a one time purchase, do ears change very much over time? (Not saying your pricing is unfair or anything, just curious.)
Are photos really enough to get a good result, though? I'd imagine e.g. internal bone structure and layout of the ear canals to have quite an impact as well.
Useful for trialing different solutions but it’s quirky.
I remember a lot of effort went into making sounds feel right in the NFS games. Including mixing engine sounds recorded from different parts of the car to match the camera's position relative to the car. We even had an article about the process of recording those sounds:
http://www.speedhunters.com/2015/10/recording-the-sounds-of-...
Speaking to the sound engineers was one of those wonderful glimpses into a whole other field that I love about working in games.
I can verify that at least parts of this list is correct. I remember getting jump scares when hearing monsters behind me in Dead Space 2, F.E.A.R., and as I mentioned earlier Left 4 Dead 2.
One thing I will say is that settings screens usually do a poor job at showing whether they're using surround sound. Games often just query the system for audio setup and show little more than whether you've selected headphones or speakers.
edit: since I thought it strange that your experience was so different from mine and that Fortnite wouldn't support it, I googled it and found an article from 2019 about how Fortnite, originally built for 5.1 surround sinds sound, now supports 7.1. My guess is there's some configuration issue with your PC.
But primarily via the analog audio outputs on the back of a PC motherboard.
If you've hooked your PC into your surround sound setup via HDMI, optical cable, or other single connection method, it is unlikely to work.
If your surround sound amp supports per-channel analog in and your PC has 5.1 or 7.1 out support, you can use 4 or 6 (5.1 / 7.1) 1/8" stereo to RCA cables to connect your PC to your surround sound system. Depending on your TV / surround sound system, you may need to mess with delay settings to get the audio to sync with the video. Having your TV and surround sound system set to "game mode" may help. Hooking your PC output directly to the TV may also help.
My understanding is using DTS, Dolby Atmos, etc to encode / decode surround sound requires paying licensing fees. These aren't paid by Microsoft as a part of Windows, by the gaming framework companies for their Windows games, etc.
I suspect Microsoft and Sony handle the fees for their consoles, but I do not know for certain. It's possible they require game makers to pay them.
Edited for clarity.
What I don't understand is why Dolby Atmos is not used much more for games. Xbox One and PC support it but only very few titles use it. PS4 and PS5 don't support it all for games (only for video content), despite all Sony's bragging about their dedication to PS5 audio. Dolby Atmos seems perfect for games, for developers because it effortlessly maps audio directly to any 3D position in space, and for users because it scales all the way from headphones to soundbars to full 7.4.2 setups.
I was royally pissed off to learn PS5 would not natively support Dolby Atmos, I have a full 7.4.1 home-theater setup with height speakers and movies and Dolby Atmos demo's sound absolutely awesome. Yet if I play games the best I can get is 7.1 which is nice, but the height speakers go totally unused. It's probably related to licensing costs, but it is extremely disappointing having waited for the PS5 for so long and not seeing any kind of upgrade to the audio.
As far as I know Dolby Atmos can be seamlessly mapped onto to 5.1 or 7.1, so from the developer perspective there should be no effort/cost to provide Dolby Atmos audio. I might be wrong about this but I assume the licensing cost would be for the playback device and not for the 'right' to bundle an Atmos audio track with your game?
There are a number of plugins like it, and I think even GarageBand has a spatial filter for it.
Was thinking through what a more modern audio game might be like given these features, it's such a rich area.
There are dozens of ostensibly high quality brands selling 500 dollar headphones with drastically different audio quality. The audiophile community would often have you believe that all such headphones are junk and a random small brand has the best headphones of all.
At least tvs have color accuracy/size/price. Headphones are not measured with any clear metric of quality as far as I've seen..
It's hard to compare when a person goes into one shop and hears an action movie, but in another one he hears a concert/symphony on a different speaker/headset.
So it ends up that most places just post the measurement curves and expect you to learn how to read and compare them yourselves.
Although, as far as I understand, headphones are both quite simple in engineering and complex with the task they try to handle.
Every person's audio perception is unique and depends on unique anatomy and unique experience (brain calculations). So it is quite difficult to normalize every data we can measure.
Still, some are trying to do it. E.g.
http://rtings.com/ - authors develop their methodology to take into account all relevant measurements, and explain their function
http://diyaudioheaven.wordpress.com/ - author does a measurement based subjective analysis of overall audio experience. Sometimes directly comparing rivals in relevant aspects.
https://www.reddit.com/r/oratory1990 - based on research from Harman R&D, measure and develops equalisation parameters for popular headphones, as well as, their weighted score (deviation from target)
The research is quite interesting in itself when addressing you statement, because it develops a statistical model of preference for audio reproduction by average human. Of course, it is not a silver bullet equal for all but many agree they enjoy headphones with "target" frequency response more.
the root of the issue is audio is highly subjective so what sounds good for you may sound awful for me. Confound that with other variables such as amps and dacs (tubes, class A, "warm" vs "cold") and it's a total mess.
Second headphones ARE measured with a clear metric, you have frequency response graphs and distortion graphs which paint a objective picture about how the headphones sound.
Mixing with such plugins is weird because you have to run everything through them and "build" any stereo effects by hand - because sounds become so localizable that they sound like someone positioned a single speaker somewhere. Just like real surround sound, it also often sounds very unnatural and you cannot really use any other reverbs or delays for sound sculpting.
As-is, the article is more like a mix of product history and sales pitch than something a hacker enjoys/expects.
Imagine an article about brtfs just mentioning when it started and that it procides snapshots. Yeah .. sure, but how? That's what interests a hacker.
Speaking of, my first time in a virtual room with the Index was super convincing because of the sound, less so my poor 970 struggling to keep up. I could move around and hear the sound sources stay fixed in a physical (virtual space). Lean towards a speaker on the desk and it gets louder exactly as you move your head. Sound is definitely underrated for creating a physical presence!
I've heard the Sony system is good, but audio has to be encoded in Sony's format.
I used to own the Smyth Realiser 8 and it was awesome. The latest version is the 16. It's rather expensive, which is maybe why the source article didn't mention it, but you can feed it almost any surround sound source. It's probably the most versatile (and expensive) of the bunch:
To use a lighting to 3.5mm connection, I think you meant the AirPods Max, not the AirPods Pro. On that note, the AirPods Max have the same spatial audio support/restrictions as the AirPods Pro.
Best case scenario I've ever had is that it's distracting as hell.