Brain signals translated into speech using AI
nature.com
nature.com
“Previous implant-based communication systems have produced about eight words a minute. The new program generates about 150 a minute, the pace of natural speech.”
https://www.nytimes.com/2019/04/24/health/artificial-speech-...
A pdf of the paper itself is linked from elsewhere in this thread.
I read your comment, went back to check comments, and couldn't believe my eyes. People think this is mind-reading.
> the article is using a benevolence trope ... to cloak the real purpose of this research — interfaces for perfectly healthy (and likely rich) people to control maybe their car or television with their minds (which would be cool!).
> Soon employers will be able to require that all employees will have their thoughts recorded while performing any work-related task, as "performance of these tasks carries no expectations of privacy."
So many people missing the fact that the technique (ECoG) requires a surgical procedure. It's not a consumer product, and it won't be cheap or easy any time soon.
Is there a Gell-Mann amnesia effect [1] for comments sections?
The employer listening to your thoughts thing? yea I doubt it. Maybe possible at some point but unlikely it would be normalized anytime soon.
There's a lot of things happened in between the pipe and outer space. Similarly, there's a lot of things happening between the arachnoid mater and the surface of the skin: innumerable liquid-solid interfaces of varying types being a major class of issues. The electric activity in the arrector pili, the galea, etc. Distance, bone, etc.
If you have prior knowledge of the person's brain activity when looking at an image, it's possible to predict (reconstruct an image of) what they're looking at.
https://journals.plos.org/ploscompbiol/article/figures?id=10...
It depends how tightly you want to define "thought", but being able to think of an object and having a computer display a reasonable representation of it seems like a good start.
It seems to me like some sort of fear of talking about commercial use cases - for example in this case the ability to talk in your mind with your phone, person-phone telepathy if you wish
In that case, 0.0001 * 1 million >> 1.0 * 100
I already have attached devices that let me enter text faster than I can think it. I think the tech only gets interesting for the handy-capable when it goes beyond words, capturing images or sounds as you imagine them, and that is a long way of.
And while I can type sentences pretty fast, sometimes I get bogged down with things like code and entering special characters and being able to subvocalize a single syllable to activate macros would be pretty nice. (Yes I know you could do that without a brain interface, but the point is that hands are limited.)
Edit: why does this comment get a downvote?
I'll blame it on the scene in II being so much more memorable. ;)
> This work was supported by grants from the NIH (DP2 OD008627 and U01 NS098971-01). E.F.C. is a New York Stem Cell Foundation-Robertson Investigator. This research was also supported by The William K. Bowes Foundation, the Howard Hughes Medical Institute, The New York Stem Cell Foundation and The Shurl and Kay Curci Foundation
(Acknowledgements section, p. 498.)
As such, most early use cases would be restricted to people with few alternatives, such as medical cases.
"This technology uses an electrode array referred to as ECoG, which needs to be surgically placed on the surface of the brain. Current technology does not make it possible to have the signal quality from non-invasive methods."
You don't think that people who have lost the ability to speak, e.g. from stroke would find it useful?
If you cure a quadriplegic, then getting approval for the general public tends to be easy (it's an easy sell) but the other way around... not so much.
Cortical implants basically form the opposite side by providing a direct interface between machines and the brain. They are already available and deployed in medical contexts. Why aren't they deployed in healthy people as well?
They have pretty cool advantages. With them, you can listen to music, etc. without being affected by outside noises (unless you want to). Imagine having a phone call in a noisy environment and you can actually hear stuff. Imagine your neighbours having a loud party but you can still sleep tight because you turned it into sleep mode where only fire alarms may disturb you. How many conflicts you could avoid!
However, cortical implants have heavy disadvantages: They give you worse hearing quality than normal hearing, they require surgery, and also need devices on the outside that are in contact with them. That's why they aren't useful for healthy humans yet. But maybe, with improvement, one day they will become.
Basically, the problems of this technology will be the same: worse quality, requirement for surgery, requirement for extra devices to carry around.
> As each participant recited hundreds of sentences, the electrodes recorded the firing patterns of neurons in the motor cortex. The researchers associated those patterns with the subtle movements of the patient’s lips, tongue, larynx and jaw that occur during natural speech. The team then translated those movements into spoken sentences.
So this means the audio example is speech synthesized from data gemerated while the person was actually reading out loud, right?
Why does the text under the headline claim "no muscle movement needed", which would imply audible speech synthesized from mere thoughts?
All neural activity (from what I saw in the paper) is neural activity while the individual is reading sentences out loud.
> Although synthesis performance for mimed speech was inferior to the performance for audible speech—which is probably due to absence of phonation signals during miming—this demonstrates that it is possible to decode important spectral features of speech that were never audibly uttered (P < 1 × 10−11 compared to chance, n = 58; Wilcoxon signedrank test) and that the decoder did not rely on auditory feedback
While the flashy part is the “speech synthesis”, the science breakthrough is actually better frames as an machine learning problem.
Imagine you record someone moving their hand across a canvas. The hand movement becomes the input, the drawing the output. The ML problem they solved is to reconstruct the output (or at least something that resembles it) from the input. In the study’s case, that’s efferent motor signals.
There is a long history of mapping these kind of signals in the sensorimotor homunculus to the respective muscles they control downstream (and some really cool stuff like prosthetic limbs can already be controlled with it), but speaking is a notoriously hard motor task and requires a lot of muscles to work in unison in very precise ways. When you implant these multi-electrode arrays, you get a few hundred more or less random single neurons, astrocytes, and local field potentials from the nether in between. Nonsense, noisy data. Being able to map this back to the result they produce in the body is technically as complex as astonishing!
I know nothing about these kinds of electrodes, but how sensitive do they have to be? Sub-microvolt? And how fast? Sub-millisecond?
Finally, on the results themselves. If I read it right, the sensors had been previously implanted into epilepsy patients, and were "re-purposed" for this study.
So I assume they patients were able to speak? If so, it's trickier to demonstrate advantages for those who have lost the ability to speak.
Hopefully there can be some further development from here into some practical applications.
Or at least reading words from your brain.