Mics that record in 3D ambisonics are the next big thing
cdm.link
cdm.link
On a side note, the upcoming 1.3 release of the Opus codec is adding support for Ambisonics-encoded surround sound.
Inverse ambisonic recording (i.e. measuring emitted sound from outside the focal point, rather than attempting to record from the focal point) is considerably easier than real ambisonic recording.
It was developed by the long defunct UK National Research Corporation that also brought us carbon fibre and the hovercraft.
It was pretty cool to rotate the soundfield through headphone playback to have the recorded audio whizzing around your head :)
I never thought at the time that VR would become popular again and provide a good use case for this technique.
for the larger capsules I had to use more parts for the frame (small pieces of bent brass) to correctly mount the capsules.
If you're interested in getting into this, you can feel good about the Zoom product. They deliver outstanding value for money. I've used them on several feature films, first as backup recorders but later as the primary audio capture platform.
Thats not true as long as
1. You record with enough near field mics
2. Your spatial resolution is fine enough to fully encompass the person's sound
As long as you do and calculate for both (or choose a hardware platform correctly) then audio processing is a simple modification within that bubble.
The problem is tracking the voice as it moves around the room, and correctly mutating the voice with the correct filters, while avoiding modifying data you don't want. My guess, is there will soon be an echo sound test to triangulate the geometry of the walls. With effective calibration and base audiographs, then much more can be done.
Of course, then its 30 seconds before someone feeds CMU sphinx each voice stream and realtime translates it into text.
I've had this workflow going on for about 4 years once I found out the XBox Kinect had 4 near field mics as well as a ir depth sensor and webcam.
If you want a jumpstart on this, install ROS, install HARK https://wp.hark.jp/faq/ and go to a used game store or pawn shop and buy a old style Kinect for $10. You'll need the USB cable so you might need to go to eBay as well.
Your workflow sounds very interesting and at this low cost I definitely want to give it a whirl, thank you!
One thing to keep in mind, is that when you plug in the Kinect, you can get the ir and webcam data trivially. However, when you load up the appropriate also commands to see the audio device, it will be markedly not present. You need to load the appropriate firmware with this tool ( https://manpages.debian.org/stretch/kinect-audio-setup/kinec... ).
Once you load the firmware, you'll see the inputs and outputs as you should. I forget the exact procedure, but I know you have to download the driver and strip the firmware.
As the ultrasonic beam travels through the air, the inherent properties of the air cause the ultrasound to change shape in a predictable way. This gives rise to frequency components in the audible band, which can be accurately predicted, and therefore precisely controlled. By generating the correct ultrasonic signal, we can create, within the air itself, any sound desired.
- presumably with any directivity desired.
https://www.holosonics.com/what-makes-a-sound-source-directi...
If practical it could keep the noise down in gaming and home theater rooms plus if the cylinder were scaled-up for auditoriums could give all in the audience front row center seats (at least acoustically).
Here's an example - https://www.youtube.com/watch?v=-p42IRDaKNc
There used to be a company selling speakers called Hypersonic or something, they were equally directional but sounded terrible.
I ask because as part of a student project we've demonstrated a delta-sigma DAC on an FPGA with very good dynamic range. We should be able to put a lot of these onto the one FPGA. It might also be extended to an ADC.
I do think that consumer mic array representation could be valuable eventually, but you’re going to need to beat the costs of ADC’s and simple multiplexed front ends. An FPGA is an expensive way to do that at volume (think ASIC).
Consider per-channel gain stages for dynamic range enhancement...
I think it will be interesting to see how this develops now Nvidia has 'affordable' cards with dedicated ray tracing.
Keywords: spatial sound rendering
Per the article:
The ambisonic mic necessitated a different method of recording the show. Instead of the standard “one-person, one-mic” studio approach that the actors recorded simultaneously, in the same room. The approach allowed for more interaction between actors, more like staging a play than recording an audiobook.
1 - https://www.theverge.com/2018/5/30/17409704/wolverine-the-lo...
If you want a 3D audio experience, it's possible to do that with just two microphones and a dummy head model [1]. Then you just play back the recording with earbuds. It's pretty fascinating that something so simple works, but you can only listen from the position where the recording was taken.
I guess if you record using one of those complicated mic arrays, it may be possible to simulate the effect of your pinnae in software, allowing you to move around in a virtual environment, and hear the 3D audio from different points?
[1] https://en.wikipedia.org/wiki/Head-related_transfer_function
>Music Research Centre's Arthur Sykes Rymer Auditorium is equiped with a sixteen speaker Ambisonic rig. This rig cosists of four high speakers, eight horizontal speakers and four speakers below the audience in the air conditioning plenum duct, which was designed to allow for this. The rig is driven by a Firewire audio interface (Focusrite Saffire Pro) which can be accessed from computers positioned in the performance area via Firewire. This rig can do up to third order horizontal with first order height.
https://www.york.ac.uk/inst/mustech/3d_audio/ambisyrk.htm
Here's some more York uni ambisonics stuff - https://www.york.ac.uk/inst/mustech/3d_audio/ambis2.htm
I'm not sure if massive speaker array systems like BEAST (http://www.beast.bham.ac.uk/about/) use ambisonic diffusion techniques, I think part of the appeal of ambisonics is the relatively small number of speakers you need for full sphere spatialization.
(It's also interesting to see this subject pop up here. Maybe because Zoom is offering a consumer option now? I've heard good things about the octomic: http://www.core-sound.com/OctoMic/1.php )
Are these companies actually understanding the science and producing well-calibrated systems that can accurately record and reproduce pressure waves in 3D? Or are they just taking a bunch of microphones and gluing them in an aesthetically interesting arrangement, and then playing them back in an ad-hoc way?
You can imagine the if you captured "pixels" of sound over a sphere the size of your head, as you turned your head, you'd be able to localize sound because mid and high frequencies would have wavelengths smaller than the sphere.
The idea of Ambisonics is to encode this spherical sound signal in the form of its spherical (spatial) Fourier transform. That way, you can drop almost all of the higher coefficients, as a form of lossy compression. In fact, most systems just capture the first-order spherical harmonics, which is enough to do a good job localizing a single sound. Conveniently, you can do this with no computation, using just 4 standard mics in the right configuration.
Spatial discrimination of multiple sound sources will be limited unless you capture higher order coefficients, which requires a much more sophisticated setup.
The second part about Ambisonics is that now that you have this handy spatial signal, you can map it back to an arbitrary speaker setup to approximate the original sound field over that sphere, if you know where the speakers are placed.
In many ways, it's far more scientific than much more common forms of multichannel audio, like stereo or 5.1. But that's kind of the downfall, because it doesn't necessarily map to that many real world use cases. In the real world, we dynamically navigate fields of sound, but Ambisonics is only going to capture sound at a predetermined point in space. If you imagine the edits in a video, the visual point of view is constantly changing. I suspect doing the same with audio wouldn't be as effective, but who knows?
If my memory is working correctly, the channels chosen were Left, Left-Center, Center, Right-Center, Right, High-Left, High-Right, Left Surround, Right Surround, Back surround, and two low frequency channels. His research indicated that the human ear has better vertical localization towards the front (potentially having evolved to detect tree-dwelling predators, for example), and experiments with dummy head recording produced inadequate results for theatrical reproduction.
Sadly it doesn't seem to have caught on (yet), probably due to the expense of having to retrofit cinemas. Anyway, it was really cool to listen to. I imagine the use of ambisonic recording rigs will greatly benefit the 360° video playback experience. Don't know what the other use cases might be yet.
https://www.dolby.com/us/en/technologies/cinema/dolby-atmos....
I have had a few going as far as denying that ambisonics could even have any kind of useful application in music at all.
The visual arts crowd seem far less dismissive, weirdly.
• it's actually about the music, not any particular "auditory experience", so technologies like ambisonics – or heck, even stereo for that matter – are perceived as gimmicky (akin to 3D TV), or
• it's about crafting a very particular auditory experience: audio is mixed in the studio so it matches exactly the artist's vision when played back to your two eardrums.
Especially the latter case I think is not served by ambisonics, because now the artist has no control over how the product is consumed. I can only imagine the difficulty involved in preparing a musical recording that sounds like quality art when the listener can – and is expected to – completely change the dynamics merely by tilting their neck. Sure you can deliver the "concert experience" but no-one goes to a concert for the ability to hear the music with their head cocked at a weird angle… so technologies like 3DIO I think already serve this purpose well.
The analogy I would put forth is a movie where the viewer can control the camera angle. (If I recall, this is actually a technology that exists to some degree, and has found application pretty much only in pornography.) Good cinematography carefully controls the viewer's attention and focus through use of set design, camera angles and focus. Once you let the viewer just look anywhere, set design becomes immensely more complicated, and the artist loses the creative control afforded by camera positioning.
Whereas, visual arts is exactly where I would expect ambisonics to find application. Visual arts is all about setting up exploratory experiences, which is impossible with traditional stereo audio recording and fixed video recording. With AR and ambisonics, now the art can exist digitally. I think ambisonics will find application in video game design too, for the same reason: now 3D soundscapes can be recorded, rather than simply synthesized.
Of course there are crossovers… I would be surprised if the ambient music crowd (think Music for Airports) does not take to ambisonics. I'm curious whether your experience differed between musical genres?
I don't think they have much use outside VR - you basically need head tracking for headphone use (or a lot of speakers)
I’ve got an H2n though mostly I’ve only been using it with line-in to record from my synth now and then, so I haven’t really used the spatial audio feature... so actually come to think of it I’m not quite sure why I am even asking :p Oh yeah, when I originally bought my H2n I did so because I wanted to get into field recording.
In the end I didn’t find many interesting things to record; traffic - boring, wind in the leaves in the forest - boring, birds singing - birds don’t sing much here, trains - marginally entertaining the first time I recorded a train coming to a stop.
I guess what I should be asking is: What sorts of sounds are you using the H3-VR to capture? And do you find any advantage in using the H3-VR instead of an H2n for said recordings?
http://www.mobilitytechzone.com/lte/news/2018/09/05/8812379....
Ossic raised $2.7m on Kickstarter to do something similar but failed to deliver, I'm hoping Creative have better engineering and deeper pockets.
This also allows us to locate sound not only spatially along the axis between the ears, but above and below as well.
I wonder if people using hearing aides, where the sound is recreated further inside the ear, are less able to locate sound sources.
https://en.wikipedia.org/wiki/Head-related_transfer_function
Short answer is that yes, the folds of cartilage (pinnae) in your outer ear are there precisely for this reason, to provide some basic directional-sensing functionality with only a single ear.
This high-end manufacturer might have a divergent perspective on those topics in comparison to some other companies. Still, they supply recent high budget productions in 3d audio, film or 3d audio sports events.