The human ear detects half a millisecond delay in sound
aalto.fi
aalto.fi
Assuming the speed of sound is around 340 m/s, in half a millisecond it travels about 17 cm, which, I would guess, is larger than the average distance between our ears.
So, using that rough estimation, I would cautiously extrapolate that we probably detect a delay under 0.5 ms, but I'd be interested to see what "detect" means exactly.
ISBN 0393090965
If you only determined direction based on relative volume between your ears, you wouldn't be able to distinguish between sounds in front vs behind you.
The relevant keywords are "Interaural Time Difference" (ITD; this phenomenon) and "Interaural Intensity (or Level) Difference" (IID/ILD; i.e., volume).
In fact, there are a few other mechanisms too. The shape of the pinna (external ear) does some filtering that allows you to distinguish sounds that produce identical ITD/IID.
The neuroscience of this is really fascinating, and the circuits have been worked out pretty well.
Good article on this:
http://alumni.media.mit.edu/~araz/sss/Sound_Localization.htm...
For anyone not familiar, here's an example of comb filtering, where reflections interfere and you get wobbles in your frequency response (shown in an fft mid video).
For example, adult mice have small heads, and ear separation is merely 5-7 mm; yet this is sufficient to locate the position of an ultrasonic squeak generated by a mouse pup at 40 kHz. This is a computational feat that requires extreme temporal precision in binaural auditory processing across comparatively noisy wetware (transduction noise, phase locking error, synaptic release noise, conduction velocity smear, dendritic integration in MSO).
And doubly impressive given that the brain has no “given” time base or oscillator to define a compute cycle. We must all build and refine our own internal set of pseudo-clocks for sensory and motor systems, in order to define the cumulative temporal context in which we are embedded.
This is crucial for the mouse to quickly avoid the talons of the owl.
More on timing in brain: @robwilliamsiii (see pinned tweet).
1. When the sound is coming from, say, the back left, the difference between the ears is far less than the width of the head.
2. It is not that the brain hears on one side "later", it is that the brain is constantly hearing two different things, what is registered on the right and what is on the left, simultaneously. Since there are often sounds coming from both directions at once, there is no easy way to match the sounds up.
Instead the subconscious has to hold each of the incoming sounds for some fractions of a second, and then combine the different sounds (with smell, etc) while matching up the different parts, using the differences to account for direction.
Since this per-force takes time, the subconscious meanwhile sends a "image" to the consciousness with the assumed incoming data based on whatever had happened before [not on what is happening now], while processing the current data to predict what image will be sent to the consciousness during the next incoming round.
This means we perceive a world created by the subconscious's guess of what should be happening now, based on past input. And the current input is held for processing. And STILL, the mouse is able to avoid the owl!
This is accomplished by the delays created between direct into the ear sound and reflections off the outer ear folds.
These reflections create "comb" filters in the audio spectrum which we learn to associate with direction. Its remarkable.
A test to prove this was so was to fill the outer ear with plasticine and perform localization tests on subjects. They could not localize sound in that condition.
The early work as I recall was at the Heinrich Herz institute in the mid 1970s.
I am suspicious that part of what this article is reporting is due to phase cancellation effects causing similar filtration that people can hear as timbre change rather than actually detecting the time delay.
(Source: My recording engineering final paper)
Have video games already emulated the spectral effects of sound direction for players using headphones, or even speaker systems with known spatial distribution? I can imagine modulating the sound coming from an in-game object to match its perceived source direction to its location relative to the player.
Edit: I'm currently working on a series of videos to explain sound direction and perception. Some of the code I wrote for my demos is or will be on GitHub.
Even a "brand" for it THX Spatial Audio
This often results in being unable to tell front/back apart, frontal sounds perceived as coming from above, etc.
What is the speed and variability of neural signals traveling through the brain?
If you get a stereo pair perfectly locked in on delay to eardrums, you can produce an extremely compelling listening experience for those in a very specific region of the room.
Finding out about all this can be revolutionary for some music enthusiasts. Once you get a good listening setup (or headphones), you start going through old things to see how the "stage presence" sounds, or if you are now able to physically place each instrument in the virtual space.
No special modern receiver in my case, just simple speakers. It seems the 2-way, 3-way speakers I grew up with kill stereo imaging. (Never mind the crossovers eat power and diminish the efficiency of the speaker — requiring a higher current amp, etc... Lovely what a small ½ Watt tube amp and a pair of full-range drivers can sound like ... and throw in a sub.)
Yeah there are ways to build crossover networks that can minimize these issues (phase shift) across the frequency range. The most ideal crossover would dissipate 100% of the undesired acoustic power as heat rather than storing it as reactive energy in inductors, but the frequency domain be a tricky beast to dance with.
The best overall approach is probably the 4th order Linkwitz-Riley filter:
https://en.wikipedia.org/wiki/Linkwitz%E2%80%93Riley_filter#...
It was magic, if you were sitting in the driver's seat. Toggling that button made the sound switch from "okay" to "magically spatial."
I'm pretty sure that lead to a conversation about how we locate sounds. Also wonder how it was implemented in 1985. I doubt there were DSP chips in there. What's the "simple" way of adding delay to some speakers?
pro audio equipment sometimes used 'bucket brigade' chips to implement a delay line (shame that you can't get them anymore)
[0] https://www.electrosmash.com/mn3007-bucket-brigade-devices [1] https://www.coolaudio.com/features-page.php?product=V3205SD
They way they work by design, clock noise needs to be filtered out of the final signal, so relatively heavy low pass filtering is standard. The result isn't very hi-fi.
The impact on the stereo field by just changing the mix of these two components is profound. No signal delay needed.
For signals that are simple L and R, sum them to get L+R and difference to get L-R. So you can use this technique on any stereo source.
[0] https://www.sweetwater.com/insync/stereo-enhancement-work-mo... [1] https://en.wikipedia.org/wiki/FM_broadcasting#Stereo_FM
But looking up photos online, I see the button I was talking about labeled as "Ambience", which kind of suggests the method you're describing.
It also turns out that using the Internet to find technical info about a 40-year-old car that was never popular to begin with is very hard!
So, psychoacoustics is incredibly complicated. There's something like 13 different mechanisms that co-operate in sound localization.
However, the bulk of it was known quite a bit before 1985, and had nothing to do with the "spatialize me" button on a specific car stereo.
There's no simple way to add a delay to some speakers unless you're working in the digital domain. In analog you have two basic choices. With passive components, you build a ladder filter, which is as the name suggests, just a long chain of low pass or all pass filters. Each "rung" only adds group delay on the scale of a couple usec, so these get very big and expensive fast. They also suffer from accumulated imprecision issues. With active components you can create a feedback loop through an op amp. This is how guitar delay pedals work, but the more delay you have the more distortion you introduce.
Technically there's a 3rd way: extremely long wires, but that's basically never practical.
Thankfully these days everything starts out in the digital domain, so you just need a controllable fifo before the DAC. Entry level home theater receivers have had this since circa 2000.
Did the math, ~170km (assuming 300,000 km/s and speed of light in copper been 90%) - that's a long wire, would also have to be superconducting or very high voltage ;).
Whatever room temperature superconductors end up costing Audiophiles will be the early adopters ;).
DSPs were a thing in the 80s, right? I guess the question is: were they so expensive that it was unlikely to find one in an OEM stereo of a mid-priced car? A sibling comment mentions that this button may have just been a stereo expando kind of thing, rather than localizing the sound stage through signal delays. I'm thinking that may be right, and that I'm misremembering the feature. It would be awesome to find some original owners manuals for the 1985 Impulse that have any mention of this feature. My DDG-fu is failing me right now.
Now, which one of those 13 mechanisms is failing on me when the damn cricket keeps "moving" around the room as I try to follow the chirp?
or an earlier one? [1]
[0]: https://i0.wp.com/www.curbsideclassic.com/wp-content/uploads...
[1]: https://www.thetruthaboutcars.com/wp-content/uploads/2014/10...
I get it was an impressive experience, but it's essentially certain it's what the other poster said: just boosting the out of phase content between the channels. This was a very in vogue effect at the time. I remember listening to the top 40 on the radio one time and Madonna had some new song where they were hyping it as surround sound and turning it into a whole event. In any case this effect can be more effective than you might assume. After all, the first consumer version of Dolby Surround was just this out of phase content run through a bandpass filter and sent to surround speakers.
I knew a family friend with the 90s version of the Impulse. Neat quirky car from what I remember. As a kid I definitely thought it was very cool.
Edit: re: extremely long wires, this is how some of the physical layer testing is done for networking equipment. Have rolls of hundreds of miles of fiber sitting on the ground to simulate large distances between switches.
Edit2: https://www.m2optics.com/products/fiber-test-boxes/multi-spo... :-)
I think a better word than "physical" domain is "mechanical" domain. Mechanical waves propagate much more slowly than electromagnetic waves.
This is very cheap and easy with opamps.
It does indeed sound magical, and makes the stereo image expand beyond the speakers.
It also does weird things to a mono mixdown and if c is too big you get a hole in the middle. But if you keep c small and use it as the final effect just before the speakers that's not a problem.
You could also use BBD[1] chips to add a ms or so of analog delay, but that's less likely because it would have been more complex and expensive.
Digital audio delays had appeared in studios by the late 1970s, but they were still more expensive than BBDs in the mid-80s, so unlikely for in-car use.
[1] Bucket Brigade Delay
It does make me curious about the recording process and effects if one microphone has lets say 50 feet of cable and another on a different musician has 100 feet of cable.
If a foot is a nanosecond for c, then 50ft difference is 50ns or about 10000x times smaller than the smallest difference. Even if the speed in a cable was 1%c, it's still under the proposed threshold.
Even so, try convincing my dad that it does not matter.
But maybe you meant that the cable lengths helped place the speakers symmetrically.
So, I presume, their listeners are hearing a timbral shift caused by the harmonics being advanced/retarded. This is not the same effect as feeding a delayed signal into one ear and a non-delayed into the other (the relative phase being used for locating a sound).
They are I guess exploring how accurately they need to reproduce an impulse response to make an accurate transducer, and from the article, I believe they have concluded that they need to be more accurate than they previously thought. Given Genelec make a range of speakers with DSP for this sort of thing, it's I guess partly a marketing campaign to convince people that there is some benefit from their DSP corrected monitors.
Of course, for a producer to have an accurate monitor is useful, but if the listening public have non-aligned drivers, the sonic benefits from worrying about this stuff are of somewhat limited value.
Most guitarists that have a rack-setup, probably have or have tried these - in practice, the effect is just a more smoothed sound.
The DRC room correction software could achieve excellent results with frequency-dependent temporal correction back in the early/mid 2000s: http://drc-fir.sourceforge.net/
If you talk into the mic with your headphones plugged into the analogue side then switch to monitoring the software instead, you can really notice the latency even though it's practically not that much. I read about a technology that's supposed to prevent angry customers from giving poor call centre employees a tirade of abuse by echoing their voices back at them on a very slight delay which is apparently quite intolerable, and I could believe it!
Sometimes I get this effect on my normal calls, and it is pretty awful. Echo wasn't nearly so bad on circuit switched calls, but now that everything is digital and has sampling and codec delays, the echos come back so much slower.
Those of you with Android devices can try this ancient web audio demo to see that the issue remains unfixed all these years later, except on certain select premium devices (certain Samsungs, etc.).
https://webaudiodemos.appspot.com/TouchPad/index.html
You should hear sound the instant a rectangle is touched/clicked and you will with a desktop browser but not with most Androids.
Apple also employs dark patterns on iOS Safari to make web games impossible:
1. No full screen. ("But the user wouldn't know how to exit full-screen - unless they buy from the App Store! We also plan to make full screen PWAs impossible because of, like, Reasons and totally not because of our 30% App Store cut!")
2. Slow WebGL and no WebGL 2.0. ("But we need to run extra shader validation security checks every single frame even though we technically only need to do it once up front! But this extra Privacy and Security[TM] is totally unnecessary if you buy from the App Store!")
3. No device motion tilt control. ("Allowing the user to consent to tilt control would violate the user's Privacy and Security[TM] - unless they buy from the App Store!")
4. No JS debug console unless you register an Apple developer account, download 10+ gig Xcode on a Mac-only computer, and ask Apple for authorization to debug javascript on your own freakin iPhone you paid for with your own damn money.
I was forced to give up using "open" web technologies for games and instead go with native code game engines (Godot, Unreal, Unity, ...).
It was like a light switch went off when you got the audio delay synced up perfectly. I guess I recall this vividly because 1) it was a huge problem and 2) the adjustment to audio delay was on the shoulder buttons, while fast forward and rewind were on the triggers. Such a large problem that the adjustment wasn’t buried in menus, but right there on the controller. Prime real estate.
If the desync goes out of these bounds then watching someone speak becomes uncomfortable.
Edit: in fact, humans are among the best animals at all at this task, I think it is not really clear why. They are on par with highly specialist auditory hunters like barn owls. May be that humans can use their higher cognitive faculties to somehow get better at the task than the average animal but it's not clear to me how that would happen.
Not coincidentally, this is exactly the type of music where lossy encodings such as MP3, Ogg or AAC fail.
So we must be talking about more subtle effects, like phase shifts, and not about how the waveform envelopes line up.
This is also helped by the instruments in the back of the orchestra favoring the low frequencies, where the listener has less access to precise timing information.
This effect is amplified even more for enclosed stadiums where there tends to be a large amount of echo, and can really be quite disorienting the first time you experience it.
https://en.wikipedia.org/wiki/Beat_(acoustics)#Binaural_beat...
I was doing just this experiment three days ago. I have a new audio interface, and used audacity to try to establish its round-trip delay. I created a rhythm track using a click sound, connected one channel's playback to an input, and recorded it. The latency was about 35.1 ms.
Audacity (the 2.x series) allows adjustment of latency compensation in whole milliseconds.
Edit: I set the latency adjustment for -35 ms so there was only the 0.1-ish residual latency.
Playing one channel of the original track and the recording together (after adjusting the channel balance for equal loudness) was quite disconcerting.
But they don't mention video matched against audo delay, that I read. If you start a song and they delay it by 10 seconds, it just starts 10 seconds later, how would you even know there's a delay? Or if there's a delay in the middle of a song, how would you know any difference, as maybe that's how the music piece is supposed to sound.
What am I not getting?
The effect is pretty wild and magical. The music, I suspect, is intentionally written to not require precise timing, and part of the charm of the piece is the musicians feeling each other out as they play. It definitely plays with your expectations of what constitutes music and how you hear timing in music.
The ICA owns the piece (a copy of the piece?), but it isn't currently on display :-(
https://www.icaboston.org/exhibitions/ragnar-kjartansson-vis...
So when my son was born profoundly deaf in one ear, as an audiophile, I soon realized "time to prove the merits of single channel."
So far that just means my everyday workstation has a single Tivoli Model One radio as the output transducer and it sounds good enough for watching films and playing Spotify on the desktop.
My intended future hifi is a single Lowther speaker cabinet driven by a modestly powered single ended tube triode amplifier.
Long live monaural!
/Acey
[1] https://www.google.de/maps/@53.5546251,9.8651002,3a,75y,192....
it feels like I could count every single echo of my steps from every single element of the fence if I could count fast enough. Similar when riding my bicycle slowly down there (rarely, because downwards, so I'm doing 50 kph there usually), ring my bell, or click my tongue.
edit: I mean, I'm hearing the difference in the delay depending on the distance.
I must be missing something, because the researcher makes this sound like a cool new technique.
Oculus, and other 3D platforms, utilize HRTF (head related transfer functions) with subtle millisecond phase shifts to localize objects.
Music in this sense is the class of all time series of combinations of such patterns that we find resonating as we are able to encode and decode emotions and feelings from these sounds (as well as appreaciation for them).
I wonder what other cool things we can do with this singal processing ability of ours as we enter the age of brain computer interfaces and psychedelics.
One possible application for something like this would be to design a loudspeaker where the time-of-flight of the soundwave travelling from a tweeter to the listener would be the same as the same sound sent from a midrange speaker. Most designs do not adjust for such differences.
The more interesting version - to me, at least - would be where you compare between the two arrivals at your left and right ear, which can do very small fractions of that even with relatively high pitched sounds allowing for sound source location based on phase.
Most designs, true. Very high-end speakers do sometimes incorporate all-pass filters in parts of their crossovers to compensate for the set-back of the apparent sound source in woofers. But very few people tilt their speakers so that the distance from the midrange and woofer to their ears is exactly the same as that from the tweeter at the listening position, and getting complicated crossovers to be reliable at 100 watt power levels is very expensive.
Time-of-flight compensation is trivial in the digital domain, so some active crossovers (which operate on line level signals, before the power amplifiers) also do this. The downside, of course, is needing one power amplifier per speaker cone, not just one per box, and the extra wires if using separate amplifier blocks rather than active speakers.
The consensus seems to be that it's not worth the trouble in ordinary listening situations.
The microphone systems mostly calibrate for frequency loss across the spectrum, time-of-flight adjustments I haven't seen but it's quite possible they're out there somewhere, but they would require a pretty nifty delay mechanism or a digitization step and then turning the signal yet again back into analog which would probably cause more problems than it solved.
https://www.itu.int/dms_pubrec/itu-r/rec/bt/R-REC-BT.1359-1-...
Btw. Drumgizmo is an amazing opensource drum plugin, that you mix like you would mix any real drum.
You want to transmit the distortions audio production made without add-ons.
https://en.wikipedia.org/wiki/Head-related_transfer_function
[0] Change the environment around the ear a bit, e.g. by putting a hand relatively close to the external ear and listen how doing that changes sounds.