“Neuroprosthesis” restores words to man with paralysis
ucsf.edu
ucsf.edu
>The participant, who asked to be referred to as BRAVO1, worked with the researchers to create a 50-word vocabulary that Chang’s team could recognize from brain activity using advanced computer algorithms.
Fifty words, that's it. For comparison, a nurse a laminated piece of A4 paper and a patient who can blink get about 15 characters per minute from the English alphabet.
Great PR, but not a substantial improvement over the BCIs from the past 15 years.
I spent a solid 6 months intimately learning the pros and cons of various existing BCI systems, as well as the exact methods researchers use to make their technology look like a breakthrough when it isn't.
Specifically, here, they've created a system that can decode one of 50 symbols - a mere 1 bit more than the English alphabet - and described the symbols as "words" so that the reader thinks each symbol carries much more information than it actually does. They've also cherry-picked the patient that responded the best.
When you peel back the hype, this is about the same performance we've been getting since '08 for an invasive, subdural, non-penetrative array.
> a nurse a laminated piece of A4 paper and a patient who can blink get about 15 characters
How is this a relevant comparison? 15 characters per minute is much less than the 15 WORDS per minute performance this work demonstrates
> I spent a solid 6 months intimately learning the pros and cons of various existing BCI systems, as well as the exact methods researchers use to make their technology look like a breakthrough when it isn't.
Then you should know this is a huge deal.
> Specifically, here, they've created a system that can decode one of 50 symbols
Previous motor and speech neuroprosthetic systems focused on pointing (controlling cursor) and more recently, handwriting (https://www.nature.com/articles/s41586-021-03506-2). This work goes gives communication rate similar to the handwriting work, but decoding speech from the motor cortex has been much less understood than that of simple motor movements such as cursor position and velocity control.
Even more impressive, the test subject is not even a native English speaker.
> They've also cherry-picked the patient that responded the best.
Bravo-1 is the only subject they had.
> When you peel back the hype, this is about the same performance we've been getting since '08 for an invasive, subdural, non-penetrative array.
This is completely, utterly false. See (https://stacks.stanford.edu/file/druid:jx921pv3255/Technical...) for survey of performance of typing BCI.
You seem to love to cite your background as a BCI PhD dropout. I should also point out that I have a completed PhD in invasive neuroprosthetics and still work in the field (not that matters when anyone can look up the sources and judge for themselves).
15 words from a set of 50 (6 bits per "word" vs. 5 per letter in the English alphabet). It's like saying a 100 baud telegraph machine can decode 300 words per second, just so long as those words come from the set of "dot" and "dash".
>but decoding speech from the motor cortex has been much less understood than that of simple motor movements such as cursor position and velocity control.
Cursor position and velocity are outputs, the input is still a self-paced motor imagery task.
>Bravo-1 is the only subject they had.
Fair cop. Maybe they'll get genuinely impressive results with their next patient.
>This is completely, utterly false. See (https://stacks.stanford.edu/file/druid:jx921pv3255/Technical...) for survey of performance of typing BCI.
I posted a link in a comment below showing an ITR of 35bpm from scalp EEG from pre-2010.
>I have a completed PhD in invasive neuroprosthetics and still work in the field
I am truly sorry for your loss.
i think it gets rather muddy as there isn't really a good metric for raw signal quality (afaik). there's cell tuning and number of spiking channels, but still not a great measure of snr for bmi work (afaik). often times people will apply measures to the outputs of their systems, like task performance, but part of the problem there is that often the state model has varying quality and suitability to task, so it can be difficult to disambiguate signal quality from state model performance.
(of course, in speech recognition they don't care, the game is to minimize WER and maximize decoding speed and whether language or acoustics (at least when they were separate) get you there, it doesn't matter)
This was something I picked up on back in 2017. I did manage to come up with a definition of SNR that made some sense (basically the euclidean distance between symbol menas, divided by the noise level along the vector connecting the two symbols, assuming the feature space was basically an N-dimensional QAM signal using features 1...N instead of amplitude and phase) - but even then that didn't take into account the fact that the noise was neither well-approximated by AWGN biased nor even constant...
And of course, as you said, you could get a bad SNR just because you're extracting the wrong features (although, to be fair, the same problem can exist in telecoms too).
Anyway, knock yourself out with research: https://scholar.google.com/scholar?hl=en&as_sdt=0%2C5&as_yhi...
Here's one from 2006 that got 35bpm without an implant (the one in the article would be closer to 60, but that's to be expected as it's invasive): https://d1wqtxts1xzle7.cloudfront.net/46069183/the_berlin_br...
Here's another from 2010 with a similar result under similar conditions: https://pure.ulster.ac.uk/ws/files/11410334/cecotti_tnsre.pd...
Within the abstract:
> The average accuracy and information transfer rate are 92.25% and 37.62 bits per minute, which is translated in the speller with an average speed of 5.51 letters per minute.
5.51 letters per minute, not words. The work you cited is not comparable to the UCSF work at all.
It seems you are measuring performance in terms of bit-rate (i.e. based on how many symbols per minute), which makes sense when you are using a cursor-based speller.
This approach is not a correct measurement of bitrate with this speech-motor decoder, however, as words are being decoded based on the syllables contained within it. The decoding model is trained to recognize 50 specific combinations of syllables, and the total number of unique single syllable phonemes is about 44.
Again, it is misleading at best to user a measure of "words per second" when you're restricted to a set of 50 of them. A keyboard that had both English and Cyrillic characters in it would have 59 unique symbols.
>It seems you are measuring performance in terms of bit-rate (i.e. based on how many symbols per minute), which makes sense when you are using a cursor-based speller.
I genuinely fail to see what the cursor has to do with anything. Communication speed is communication speed.
>This approach is not a correct measurement of bitrate with this speech-motor decoder, however, as words are being decoded based on the syllables contained within it.
Just as symbols in cursor tasks are decoded based on relative position?
^by convention
Detecting (in)activity in M1 during imagined movement tasks has been the standard decoding method for 30+ years.
>This new process lends itself to huge amounts of training and machine learning to vastly expand the vocabulary.
Not without degrading accuracy. The fundamental limit to these systems is the human operator; not the decoding algorithms. ML has advanced dramatically since the turn of the century but almost all BCI performance gains have been from new sensors.
This invokes the feeling that it'll be impossible to keep computers out of our bodies at some point, which is scary on so many levels.
But like any tool, if I look at all the good that can come out of it - simply amazing things are possible.
Imagine that with a brain interface...
I could imagine easily going insane if I were forced to interact through a piece of technology that would not do what I wanted. It's bad enough when my phone won't do what I want.
Context helps - if it looks like a duck, quacks like a duck etc...
Edit: this is a possible alternative meaning for "Duck Typing"...
But to your point, would the system understand thought?
Knowing the difference between:
> "I'm sending two to you" or "I'm sending to Two One Smith Lane."
requires a lot of knowledge/intent if indeed you can't think in terms of individual characters.Systems like the palm pilot graffiti interface, or some optimized symbolic gesture system will become prevalent, but imagining typing will probably carry us sufficiently until better systems are worked out.
Is it the typing speed that limits the amount and quality of code one can write in a day? To me, it's mostly finding out things about the subject area, finding the right APIs, code search.
Of course, when typing is fast enough so that you never break your flow, it is important.
The protocol for sensory input has likewise been prosthesised by Paul Nach-y-Rita’s tongue grid array for vision.
https://www.researchgate.net/scientific-contributions/Paul-B...
Actually we have some idea and one theory is that they are not as distinct systems as we’d like to think, at least they work bidirectionally and reciprocally: language structures back the thoughts themselves, so there is a way that externalization actually feeds “inwards”. See distributed/embodied cognition.
As for the popularity of linguistic determinism, who knows. I wouldn't trust a linguist who took the strong Sapir-Wharf hypothesis seriously.
What’s left is all the logic between input and output.
We gain a pathway for investigation backwards. What neural substrate is active in generating the speech data? Work backwards from there.
We already have a useful pathway for investigation: data from speech/sign of a speaker of a particular language. In fact this already yields a kind of overabundance of data. The tricky part is finding the right kind of data via careful experimentation and organising it, by building explanatory theories of language that meet the conditions of evolvability and learnability.
(Which is not to say that cognitive sciences can't shed any light on this, but studying the brain independently of linguistic experimentation is a dead end IMO. Brain imaging studies coupled with linguistic studies have yielded interesting results - see Andrea Moro's work on "impossible languages").
More importantly, the fact that any function can be represented as a NN does NOT mean that any function can be learned through known NN training mechanisms, not even in principle (i.e. given finitely arbitrarily many examples and unbounded but finite time).
And of course, there is always still the vague possibility (which I don't personally subscribe to) that the actual function is not Turing computable, which would mean it certainly can't be approximated by an NN.
There has never been a dearth of universal function approximators, polynomials can do it, splines can do it, sine/cosines can do it. Being a universal approximator is hardly unique or special.
There is absolutely something special about DNNs but being universal approximators is not one of them.
Being able to learn a function from data is very different from being able to represent that function.
Being a universal function approximator doesn't magically solve every problem.
Big tech isn’t far away from this. Scary to think what kind of applications can be derived from this technology. Imagine a hi-tech polygraph that maps your brain activity to speech—-do you have plausible deniability if something incriminating blurts out?
I don't think that giving everyone the gift of good grammar would create a world of geniuses but it might do what every other useful tool in history has done, help "level up" those at the bottom to a better standard (of, in this case, reasoning) and free those at the upper levels to really create.
Separate to idiom is the process of rewriting, whereby rough thoughts are honed to sharp points.
“I have rewritten — often several times — every word I have ever published. My pencils outlast their erasers.” ― Vladimir Nabokov
“Revision means throwing out the boring crap and making what’s left sound natural.” ― Laurie Halse Anderson
“Secure writers don't sell first drafts. They patiently rewrite until the script is as director-ready, as actor-ready as possible. Unfinished work invites tampering, while polished, mature work seals its integrity.” ― Robert McKee
“When asked about rewriting, Ernest Hemingway said that he rewrote the ending to A Farewell to Arms thirty-nine times before he was satisfied. Vladimir Nabokov wrote that spontaneous eloquence seemed like a miracle and that he rewrote every word he ever published, and often several times. And Mark Strand, former poet laureate, says that each of his poems sometimes goes through forty to fifty drafts before it is finished.” ― Susan M. Tiberghien, One Year to a Writing Life: Twelve Lessons to Deepen Every Writer's Art and Craft
“I do so much writing. But so much of it never goes anywhere, never sees any light of day. I suppose that's like gardening in the basement. I don't publish so much of what I write. I just seem to plow it back into the soil of what I write after it, rewriting and rewriting, thinking that somehow it gets better after the fifty-second-time around. I need to learn to abandon my writing. To let go of it. Dispose of it, like tissue.” ― J.R. Tompkins
“Writing a first draft is like groping one's way into a dark room, or overhearing a faint conversation, or telling a joke whose punchline you've forgotten. As someone said, one writes mainly to rewrite, for rewriting and revising are how one's mind comes to inhabit the material fully.” ― Ted Solotaroff
Or conversing with an American, the spelling might tip me over the edge! ;-)
One of the most brilliant inventors I know has terrible written grammar due to dyslexia. I'd be careful with such blanket statements.
something's gotta give. and im pretty sure it'll be your privacy
Surely brain computers will be a pretty big investment for people; would embedding advertisements be worth it for companies? I'd imagine most people are willing to pay another few thousand dollars on top of the very high cost to avoid any sort of obnoxious features, for both personal and potentially socially-influenced reasons.
It's hard for me to see how anything could be worth more to companies than straight profit.
But what I was saying is that even if you could pay for a brain computer, many less fortunate wouldn't be able to, which opens the road for advertising subsidized (surveillance) brain computers.
you think fb will make people pay for it to receive like notifications in your brain?
Maybe we already are all just computers in a simulation already. Ghost in the shell thoughts are rapidly coming back.
My take (apologies in advance for the self-indulgence): A commonly-claimed revelation or drug-assisted-insight is that "we are all one consciousness experiencing itself subjectively, there is no such thing as death, life is only a dream, and we are the imagination of ourselves". This is a foundational concept in the idea that "cybernetically-assisted individuals" is only a minor extension of our current reality, the tiniest blip upon our shared planetary history's march of progress. It should be treated as such, but hopefully better managed than our previous technological accelerants like fossil fuels...
for example,
and those 18 words can turn into a lot more
Is 74% decoding accuracy good in a sentence of a very small length? do we know if the misses are at least semantically close or random?
It's very weird and suspicious to me that they mention the Up-To performance in the video, without any explanation or statistical quantification or context information. Sounds like they really wanna make this sound better than it is, in this video at least. (the info is likely in the paper)
Imagine poorer people will be forced to wear Facebook Talk device, that will listen to your thoughts and then project adverts directly into your eyes. They will also know when your disability check is coming and what are your desires and needs. They'll for sure show you something tempting to buy...
My first thoughts as well.
While this iteration requires implants and a conscious effort to function, I bet a year's-worth of salary that future iterations will be able to function without either.
I personally think thought crime is an inevitability at this point. I feel sorry for future generations.
The Dobelle Eye for example. In the early 2000s Dobelle Eye brain implant allowed blind man to drive a car in a parking lot using cameras that feed video into his into the visual cortex. They had to make the operation outside the US (Lisbon) and patients had occasional seizures even then.
Wired article: https://www.wired.com/2002/09/vision/
I guess if they are encoding and decoding voice from these...does it prove direct neural encoding/decoding from LFP/single units, if they can generate robot voice from someone who can't generate sound and thus vibrations on their own?
Most previous brain signal translation has been adaptions of typing in some form. This is directly translating signals intended for vocal cords, IE for vocal speech, into text.
https://www.brain-injury-law-center.com/blog/brain-activity-...
A complete lack of brain activity is brain death. Brain activity is often slowed in comatose patients, but that varies from patient to patient.
https://www.mayoclinic.org/diseases-conditions/traumatic-bra...
Human speech is generally about 150 wpm in english. I've read that the informational output is relatively consistent between languages, so in a language like german you get bigger words and fewer wpm. Assuming that's correct, 150 english wpm is probably close to the processing speed of the brain, minus the overhead of converting thoughts to speech.