> The headphones send that signal to an on-board embedded computer, where the team’s machine learning software learns the desired speaker’s vocal patterns
Their "AI" is good ol dumb machine learning
If you have good eyetracking, a microphone array and decent object tracking on your AR glasses, then you don't really need much "AI" (ie you have access to https://facebookresearch.github.io/projectaria_tools/docs/AR...)
but its not quite possible to do it all on device yet. However its not far off.
Edit: I found the link to the paper. It isn't stream splitting so much as it is GPT-assisted beamforming estimation. Good stuff for sure.
But this is super expensive since you need calibrated mics etc.
The biggest advantage of neural nets in this field is that you can use a dirt cheap microphone and postprocess it so good that it is good enough or even very good for humans.