This is what computers "hear": http://www.ling.ed.ac.uk/images/waveform.gif - in fact, this is ALL they can hear. When you hear 44.1khz, audio, that is 44100 bytes of information (a value between -128 to +127 per byte) per second. There is no hidden metadata behind the waveform. At 100% zoom-level, the waveform IS the audio. Theoretically, you can take a screenshot of a 100% zoomed waveform and convert that to actual sound with absolutely no loss of data (of course, you'd need a really high resolution to show a graph 44100x256 pixels in size. Now given such a waveform, how would you convert that to plain-text?
As an example, try recording "ships" and "chips" into a mic and view the waveform. See if there are any patterns you can identify between the two waveforms. I've done it over a hundred times. There isn't an easy way to discern if the letter was "sh" or "ch". Yet our brain does it so very easily thousands of times every day. So, failing easy pattern recognition, we have to use frequency analysis, DFT, and tons of AI.
My unscientific gut-feeling is that we're going about all of this the wrong way. We are using the wrong tools. Discrete, digital computers will never be able to tackle problems like pattern recognition in their current state. Switching to analog isn't going to improve anything either. I don't know what the correct instruments/devices will be but I know programming them will be very different.