I guess ultimately you're limited by your interface to the brain (something maybe neuralink could address?), but even still shouldn't there be some "magic" machine learning stuff going on in this space? Not at all an expert in audio processing, but I would imagine that we could do a decent job distinguishing voices from background noise, automatically thresholding sound in different environments, and possibly modulating sounds into different frequency bands when there are e.g. multiple speakers.
For all I know, stuff like this is already going on.