Generating natural-sounding synthetic speech using brain activity
humanbioscience.org
humanbioscience.org
> The researchers also found that the neural code for vocal movements partially overlapped across participants, and that one research subject’s vocal tract simulation could be adapted to respond to the neural instructions recorded from another participant’s brain. Together, these findings suggest that individuals with speech loss due to neurological impairment may be able to learn to control a speech prosthesis modeled on the voice of someone with intact speech.
As per paralyzed individuals that is a primary target of this area of research. Things get considerably more complex in those cases however as they would need to start out with a pre-trained model which couldn't be naturally adapted by listening to their own speech. Additionally any brain damage which many have contributed to their condition can impair some of the signals in question. Overall the particular problem seems to be advancing from when I worked with a lock-in patient last, but there's still a good ways to go.
> Even when the researchers provided the algorithm with brain activity data recorded while one participant merely mouthed sentences without sound, the system was still able to produce intelligible synthetic versions of the mimed sentences in the speaker’s voice.
That already seems borderline usable for people with speech. I wonder if it could be made to work when you're merely thinking about the mouth movements without actually making them. That would be ideal.
It might even be the killer app to make elective brain implants mainstream. (Of course if this could be done without hazardous and expensive implants, all the better.)
If the software assistants are sufficiently useful and tightly coupled with the human mind, I think it quite likely that the line between self and software might get blurry for some users. The ability to think a question and hear a correct answer as a voice in your head is the sort of profoundly powerful user experience that I think might plausibly alter the assumptions people make about what it means to be themselves.
If these software assistants become apart of the users' own mind in their own perception of themselves, what responsibilities do the owners/operators of those systems have to their users?
I guess we'll cross that bridge when we get their, but the relative immaturity of FOSS software assistants is starting to unnerve me. In 2040 when Amazon starts selling "god in a box" to the general public, a two-way telepathic connection to a state of the art quasi-AGI living in the cloud, will there be a viable FOSS alternative?
Edit: strange I didn't think about it directly, but that is the Borg from Star Trek.
Toddlers are, today, growing up with iPads, YouTube, and Alexa/Siri as just a natural part of the world. As far as I've observed—admittedly not much—educators and parents are far behind, I would speculate both because grasping the indirection-of-agency that this sort of technology creates is a heavy abstract task of the kind that doesn't seem to filter through those parts of civilization well and because the technology has the ability to change too quickly to counter any attempts to pin it down. And pinning it down in too static a fashion could have its own horrible effects.
“FOSS” in the original sense is largely a distraction here (even though it is still an important idea), given that we've wound up in the “programming is specialized” world. The dynamic characteristics involve how agency flows through systems, and in the presence of highly distributed and often SaaSS (Service as a Software Substitute, as the FSF describes it) systems, being able to alter the source code isn't a solid defense even at “skilled programmer” speeds: going against a rushing current just means you get torn apart as soon as you touch the world. I think we need a new word for what you probably meant but which I don't know how to articulate well.
I guess the tighter coupling than what it is now is what makes the idea repulsive.
Also, as far as FOSS alternative goes, I wouldn't count on it. It's not so much the code - in time it'd be the huge data-crunching that counts, something only big corporations are able to do.
But then again, the thought of a laser cutting your cornea was probably incomprehensible too at its inception.
they can produce words from brains even when the subject does not speak them. interrogation tool?
What I mean is, when we think about moving our arms the same parts of our brains activate as if we are moving our arms. Maybe, when we think about talking a similar thing happens.
http://webcache.googleusercontent.com/search?q=cache:https:/...
Copy the link above to view the research
Anyway, putting electrodes inside the brain is not for common public, is this at all possible without those intrusions?
Disclaimer: I know very little about how to actually do ML stuff, it just seemed like something that'd be possible in the near future.
Style transfer is trickier to do with speech than images! One significant issue is the lack of a good "content" versus "style" distinction. In images you can get great results by calling the higher-level features of an object classifier network "content" and holding that constant. Some people have tried this for audio with e.g. a phoneme classifier, but there are additional characteristics (such as inflection) that relate the emotional content of speech which wouldn't be held constant.
Another issue is that much of the speech classification work is done in spectrogram (or with further processing MFCC) space, which lets you treat audio similar to images and leverage a bunch of technology that we have for classifying those. But for synthesizing speech, spectrograms aren't a fantastic representation, because small errors in spectrogram space can translate into large errors in the waveform which are very clearly audible, and humans in general are pretty sensitive to audio errors. There are cool neural spectrogram inversion methods out there which can help, but those should still be trained to be robust to the kinds of errors that a style transfer algorithm would make, so it's still pretty tricky.
My company, Modulate, is building speech style transfer tech; and we've found a lot more success with adversarial methods on raw audio synthesis, where the adversary forces the generator to produce plausible speech from the target speaker!
One of the coolest parts of the kind of BMI research in this article, to me, is the potential to buy back some latency margin for speech conversion! If you're working on already-produced speech, there are super tight latency requirements if you want to hear your own speech in the converted voice - over 20-30ms for the entire audio loop, and you start to get echo-like feedback that makes speaking difficult. Even without looping back, you don't want more than 100-200ms of latency in a conversation before it starts impeding the flow of dialogue. This means your style transfer algorithm gets almost no future context, and limits the kinds of manipulations that you can do (not to mention the size of the network that you can do them with, depending on available compute power!).
For actual patients using the system however I'd expect it would be beneficial to keep the output as low latency and as unmodified as possible as this will help the individual learn to control the system as nerual activity shifts over the course of time.
[1] https://www.uni-bremen.de/en/csl/
edit: small correction of my bad english :P
either way cheers for getting something to work well!