Live-caption glasses let deaf people read conversations [video]
youtube.com
youtube.com
There are a few issues that are unanswered by this video (which isn't intended to be a technical deep dive, but I don't see any related links in the video description):
1. How do these glasses handle multiple simultaneous speakers? Based on the display I saw, it shows the speakers' words sequentially, which starts to fall apart in real-world environments, especially group conversations. This is a big problem, and wider adoption is contingent on handling this elegantly.
2. These appear to be the classic "smart glasses" display style that's pervasive in consumer head-worn displays today, where content is projected at a fixed depth in front of the wearer. Because the captions aren't anchored at the same focal distance as the speaker, the wearer's eyes will swap between the captions and the speaker's faces, which is a tiring activity, and can make the wearer feel like they're not part of the conversation or being rude.
3. As mentioned by another commenter, this is a useful idea for people who lose their hearing later in life. That said, this is less (although certainly still) useful for people who have congenital hearing loss and primarily communicate via ASL.
All in all, it's exciting to see growing interest in this space, as it's easily extendable to people learning a new language or navigating a foreign country. I think offloading the speech-to-text to a tethered mobile device is a good choice (though it would be nice to do low-latency wireless transmission).
As usual, this marketing seems most directed at normate ideas about what disabled people want / need, but the tech seems very cool and there does seem to be potential. Without looking into the product deeply it seems like there are D/deaf people on the team, which gives me hope.
I do wish that we would just embrace the idea that using machines to make information available in many mediums is something all people can use and appreciate.
> "This is a classic curbcut in the sense that it will help those with [edit: hearing] as much (if not more than)..."
1. For multiple simultaneous speakers of comparable volume, it’s only as good as the underlying speech-to-text engines we’ve implemented/integrated, which is currently not very good. It’s active area of research and engineering for us and we believe we’ll make strides to improve things; but, as you rightly point out, solving the crosstalk problem is very difficult. For the more general so-called "cocktail party problem", we can do a good job of filtering out more distant/lower volume voices and other environmental noise. Choosing the right microphone can improve things further, for example by pairing a noise canceling Bluetooth lapel mic.
2. We allow one to project the subtitles at varying depth, within the capabilities of glasses. We're seeing an effective focal depth range for fixed apparent size of about 0.5m to 3m. If one also allows change in apparent size, to simulate perspective scaling, the range is higher.
The thing you could do is train a localizer on the separated audio. Phase is estimated by the source separation process, so you can actually train an ML model provided you have some ground truth (e.g. estimated human locations from camera detections)
I haven't seen it provided as a commercial service or free model yet, but there is open source code for Mixit that lets you train using the open source / canned FUSS dataset.
https://github.com/google-research/sound-separation/blob/mas...
Re: #2. I’m assuming the varying depth is manually-controlled? Or is it automated by some method? If it’s manual, can the adjustment be made while transcription is active? In other words, can I change the focal distance to match the speaker without interrupting the speaker?
All in all, cool stuff! Best of luck with the work.
Indeed, as Dan said, one can change the depth on the fly. In fact, one goal of the development team is to make as much functionality as possible changeable on the fly. For example, you can currently change subtitle depth, pinning, and size on the fly, spoken language, subtitle language, microphones, and audio settings on the fly, etc.
I’d love for every setting and feature to support on the fly changes. That said, some things are currently fixed for a session, such as recording audio, and some third party software we utilize is less dynamic and forgiving of changes on the fly. For better or worse, in our software world of today, the sage advice of The IT Crowd “Have you tried turning on and off again?” still seems to hold with pragmatic force. And it still holds with XRAI ... sometimes ;-)
If you had more microphones placed in multiple spots on the glasses, such as up to 5 microphones - 1 in the center, 2 on the frames/end pieces, and perhaps a final 2 on the arms/temples. Then that would be able to catch conversation coming at a person from behind them, the sides, or directly in front, etc.
This would be such a cool use case for the latest ChatGPT tech when it gets faster in the near future.
Could you have an AI model that extracts some characteristics of the speaker's voice for each individual word, then translates that to color and font?
If the model was not confident about a word it could show slightly blurred, if it was loud it could be bold, perhaps(Although there's some stereotype issues) you could use different fonts for different pitches, whispers could be grey, quiet could be transparent.
Maybe there's a language model that can pick up overlapping words if you don't have the constraint of needing to sort them out into who said it, just show all the possibile words that could have been said by anyone stacked together, in a "not sure" color, and maybe the wearer would eventually learn to figure it out without much effort?
You could also try to stay consistent so the same speaker gets the same colors I'd possible, and also not reuse colors for new speakers that have been recently used, to best make use of the limited bits of data in font and color.
Maybe just by showing all the words from every speaker all together like that, the wearer would be able to figure it out even if it made mistakes in the speaker identification?
As for the rest, thank you for wonderful ideas! Everything you propose is technically possible. The difficulties arise first in assessing the increased benefit to users versus the increased complexity of user experience, and second in prioritizing the work versus other features. Over time we do hope to add additional selectable “skins”, which is to say different UI designs, that allow users to choose UI’s from a wide range. Everything from the simple to complex layouts, from accessible to exotic color palettes, from professional to playful themes, etc. I could definitely see more advanced visual representations of transcription uncertainty showing up in such optional skins.
She has a different "speech mode" than I: she speaks while I'm reading or washing the dishes or whatever, and sometimes it's to herself, sometimes to Siri, and sometimes it's something she wants me to know.
I noticed that my wife stopped using my name when it became more ambiguous as to who she was talking to because our kids are now older and conversations (really instructions and queries) are at the "adult content" level rather then "child content" level.
Retraining her will be challenging.
Someone primarily communicates with ASL and then there's me that doesn't know ASL. I can speak to them, and they can read what I've spoken. That works pretty well. They communicate with me via text to speech, or (I guess in the near future) ASL to speech - however that will work.
I mean, that's awesome.
Which.. aside from that, you could do this anyways. When I first started living with a deaf person, I just wrote things down on paper, and they wrote back. The glasses have some benefits over paper, but significant drawbacks as well, and many deaf people I've met have been trained to read lips.
This isn't adding much.
We can communicate without seeing each other.
It’s a decently-common occurrence. There’s a significant amount of literature pointing to the fact that people who are DHH avoid group conversations because they can’t follow along (see references). This sort of falls hand-in-hand with speech localization, which is another unsolved problem (and likely is unsolvable without true AR).
https://journals.sagepub.com/doi/10.1177/1084713807301322
The last time I worked on anything here, there were a number of problems with transmitting highly compressed, low resolution video data. The consumer devices could not handle just sending the packets. In my project we were annotating real time video and only sending back the annotations, but even that would cause devices to overheat and the applications to fail in really interesting ways.
My take is that this is a huge UI improvement for AI speech to text, which a lot of Deaf people are already using to listen to conversations. It seems particularity great because it allows this technology to provide situational awareness while, for example, walking.
It's important to remember, though that, for conversations where you're trying to include a Deaf person who isn't good at speaking or chooses not to speak, speech to text is a fundamentally unequal communication modality. They will be able to "receive" but they won't be able to "transmit", which makes for extremely lopsided conversations. There is no substitute for taking the time to learn a sign language or to have conversations via writing (no substitute for sign language as it requires a lot of patience from both parties).
Not every deaf person uses sign language and there are many different sign languages in the world. American Sign Language(ASL) is but one of these languages.
Specifically, sign languages are not visual representations of existing languages (e.g. ASL and English) but completely different languages altogether.
I'm sure other signed languages have rough equivalents in their regional areas as well. I know Mexico has a few different signed languages, though I only have passing familiarity with 1 of their signed languages and it's definitely not a representation of Spanish.
Speech to text was "solved" a long time ago but I've seen it take many years to become as usable as it has recently. And it still regularly is frustrating to use for me!
Let's take an example: the sign for "send email" in ASL (http://www.lifeprint.com/asl101/pages-signs/e/email.htm). If I point at you at the end of the sign it could mean "I will send you an email." If I point at myself it could mean "You should send me an email" or "Did you send me an email" depending on my facial expression. If I point off into space it could mean "I'm sending an email." If I start by pointing at you and then end the sign by pointing off into space it could mean "You should send an email." So your translator AI needs not only to understand the facial expressions and movements of the signer, but also the spacial relationships of everyone in the conversation. And that is just one aspect of the difficulty - there are many other features of sign languages that are just as hard to translate.
Perhaps this is the sort of thing that future AI systems could do. But it is quite complex.
Hard of hearing people have the same problem in person. We aren't really at a place yet where someone wearing hearing aids still has 3d hearing. So like on a conference call, they can't figure out what's going on when four people are talking at once.
I know a partially deaf kid who prefers socializing on Discord, because of this. Everyone is equal and all conversations have to be 2 people at a time.
1) The cocktail party problem is still a WIP. This is a very hard problem to solve.
2) These are not 'viewer' glasses they are 3DOF glasses which support moving and pinning the subtitles in 3D space
3) Whilst we targeted the Deaf and HOF to begin, we see broad applicability beyond this
4) You don't need glasses to test it, just an Android 12+ phone. Download and try it. We'd love your feedback https://link.xrai.glass/app
> 4) You don't need glasses to test it, just an Android 12+ phone.
This exactly pointed her. Some deaf people will use any dictation software on the phone, and look at the phone when needed. This glasses instead will cover some field of view. Note that view for deaf people is more important and used than for the rest of us. She couldn't see the improvement of using bulky glasses instead of lowering your eyes to the phone.
Personally I think the endeavor is admirable and wish you best of luck. Also, as other comments say, I think this product might be more desirable for HOF and late in life hearing loss sectors than born deaf people.
This is cool tech, that could be used to help people, but it comes with lots of potential for new forms of evil that were not possible without it. Considering that I can't remember the last time I bought a product using a new technology that wasn't also designed to work against my interests, I'm immediately skeptical of any device that can't be used offline and especially one that requires being connected to cell phone apps.
The answers to this determine everything. Treating privacy-first as the moral and ethical default we should expect everyone to start from is a wonderful idea, rooted in compassion, kindness, and a foundational respect for human rights. It has also been an abject failure to date.
We should not expect the future to be different unless we are willing to be realistic about the economics at work. Otherwise the market gap will remain in the realm of the wonderfully hypothetical forever.
the fact is that no amount of money a customer can pay will ever be worth more than taking that customer's money and then also selling every scrap of data that your product can get its hands on on top of it.
We need to stop accepting "But I can make more money by screwing you" as a valid excuse for how things are. It's true, but still not okay. If the economics are always going to favor exploiting people, than I guess if we don't want to be exploited we need some very powerful regulations and all of the oversight and enforcement that requires in place to change that situation. I'd prefer that to blind faith and optimism. Until then, there are a number of things we should be insisting on in products as consumers to help protect ourselves. "works offline" is a really good one in a lot of situations.
https://github.com/TeamOpenSmartGlasses/OpenSourceSmartGlass...
ASR is done locally on the user's phone.
I suppose it would be a nice feature if they saved all of your conversations for later? The translations are too imperfect for any legal matter use.
the problem isn't that it's worse than a cell phone. The problem is that it isn't any better. All of the privacy and security problems that exist with cell phones now exist with this product. You're locked in to using a cell phone which, for most people, is a hardware platform that Google controls and exploits at your expense.
You can't secure your cell phone and keep it private, and so nothing on your cell phone should be assumed to be secure or private.
If I'm in my doctors office, or working with a client, or doing anything where I don't want Google and who knows who else listening I can turn off my phone, or put it away, or leave it at home, and be reasonably sure that I'm not being eavesdropped on (there's always some risk of three letter agencies listening, but you can't do anything about that), but that's not an option if I need my phone to transcribe every word being said to me and even in situations where I currently leave my phone on it isn't normally sending every word it hears to the cloud to transcribe and send back to me either.
I like that this can work with a cell phone app, but it'd be much better if it could be used entirely offline, and ideally on different hardware as well, including a PC.
Imagine being able to run something like this entirely offline connected to something the size of an ipod or even a nano running linux. Imagine being able to build your own tools to interface with it. Add your own substitutions/annotations to real time text. Add dictionaries, or translation features.
This could be really cool even life-changing tech for people, or it could just be one more technology that's convenient but ultimately used to exploit people.
There won't be any ads, because people would just use their phones instead which don't have ads. Specifically Google Live Transcribe, Otter and the like. Those require a data connection to the network, but there are versions that don't need the network at all. E.g. Chrome's Live Caption option. Eventually as technology becomes more power efficient and miniaturized it won't need to be paired to a phone.
The advantage of glasses is that people find it very distracting seeing a phone scrolling away, my GP stares at the phone instead of me because he is fascinated by it. Sometimes you can't be holding a phone up if something is being worked on. The glasses would also allow for a bit more directionality. It's a promising tool depending on how well it is implemented.
I hadn't thought that you would be, but I don't know anyone who hasn't had at least one service they used change their terms for the worse over time, especially once it becomes successful or ownership/management changes. Better to safeguard against such problems before they are problems than come to depend on a device like this only to find yourself stuck when the rug gets pulled out from under you.
> We are soon to release purely on-device transcription
This is really great! It's so important that people don't have to worry that things said between them and a lover, or a doctor, or a client, or a therapist is going to end up exposed to anyone else.
If you really want this device to be able to help people without opening them up to exploitation, please try to develop it to be as open as possible. Making it easy for developers to use it with whatever other hardware they like (a PC vs a cell phone), or even allowing them to extend it by adding new or customized functionality would be ideal and could help people in ways you hadn't even considered.
This is a cool product and a great use of the technology, and I'd love to be able to recommend it without reservation.
I doubt it. This isn't the kind of cheap mass-market device where running ads is going to make you a big profit.
This product is already targeting some specific demographics. I'm sure a lot of people would be willing to fork over money for access to those eyes just as I'm sure a lot of people will want access to what's being said or looked at. It'd be very nice if we didn't have to worry about that sort of thing at all, but here we are.
Sure it won’t solve the issues faced by the deaf community but that’s only a tiny portion of the people handicapped by difficulty hearing.
We just got the Nreal/Xrai setup a few days ago for deaf from birth wife (hearing husband) She grew up lipreading but integrated more with signing and deaf community as an adult. She has a cochlear implant but can not understand language from sound alone. And really doesn't enjoy hearing that much unless we are watching a movie etc where the sound is 100% linked to the visual.
Initial reaction to the setup is. 1. Impressed, hopeful, excited 2. A bit complicated technically. More stuff to deal with. not an everyday thing. 3. Phone battery usage high. Maybe 3 - 4 hours 4. In the right situation they will be really powerful. 5. Need more control over the interface eg show/hide 'listening' icon etc. Can be distracting. Move subtitle position (maybe you can) 6. Processing delay can make you more an observer of the conversation. Response time is delayed enough to interrupt the flow a conversation. (satellite tv connection interviews)
The number one barrier to using them is having everything ready for the moment they are needed. You need to plan ahead. Takes a few mins to set up.
All the other high end ideas can be set aside while the core function is dialed in.
We really appreciate the effort and hope to contribute.
I can imagine users going, "I'm deaf, not blind."
Magic Leap has screens that can adjust opacity at the pixel-level.
I currently use googlemeet and recently switched to Otter.ai for recording/transcription. Unlike the previous transcription tools for journalists that I used in the past - Otter.ai generates the text live on the laptop screen while we are talking and even corrects itself to make sense as the speaker reveals contextual clues.
It is a huge help, and I had wished there was something for real life conversations like this is for screen conversation.
Good news for deaf people. You only have to watch newscasters speaking through a pasted-on permanent grin - as if every word is ee - to know that lip reading is garbage.
¹ Afaik this applies to all other sign languages outside english too. Signed Exact English exists and probably other-language equivalents too but I've never met a native speaker.
https://www.lifeprint.com/asl101/fingerspelling/fingerspelli...
Native ASL speakers who are completely illiterate in english certainly exist, and I'm not sure at all if they know or use finger spelling.
For example, it would be like me saying, "'h', 'e', 'l', 'l', 'o'" instead of whatever translated 'hello' in ASL.
Here's a fun example. ASL allows, maybe even requires, negation after the statement. An interpreter friend of mine was interpreting Wayne's World in a mixed crowd. The whole "<statement>... NOT!" joke gets laughs from the hearing audience and the Deaf audience doesn't understand why.
However, ASL also makes much, much heavier use of rhetorical questions than English does. You might even introduce yourself with "MY NAME WHAT? [NAME]" (i.e., "What is my name? [Name]"). So perhaps it would just look like you're doing that.
(Disclaimer: I don't know ASL. I know some Irish Sign Language, which is related, but dropped out before completing my interpreter training. I have a bit of a fascination with sign language linguistics, but I'm no expert.)
It's true. 70% of deaf children learn to read English [1]. But I don't think that's enough to consider it a safe assumption.
[1] https://www.handsandvoices.org/articles/education/advocacy/w...
There's "cognitive impairment" and then there's "nobody bothered to communicate enough for a kid to build communication skills."
I can see at some point here being able to wear AR glasses that overlay hand signing over the speaker
It's the same language you already know, but you're missing out on one of the primary ways people use it.
For example the sign languages spoken in the US and in the UK have different ancestries and are not mutually comprehensible, despite both countries using english as spoken languages.
Sign languages have multiple articulators: two hands, face, eye gaze direction, shoulders, trunk. These can all work together to show multiple things at the same time. Spoken languages can really do only one thing at a time (with a few minor suprasegmentals such as tone).
You can construct a signed version of a spoken language, which may be useful for things like quoting book titles and other cases where you need to represent the exact words of a spoken language in signed form, but it's not common to use that for everyday communication, because the hands move a lot slower than the small muscles of the mouth and throat.
(Linguistics is a fascination of mine. Sign language linguistics are especially interesting.)
Well, yeah. I guess my point is that ASL is pretty much a foreign language.
I'd compare it to the situation with the english language worldwide - since english is the lingua franca, so to speak, many countries around the world teach it as a second language. If you don't learn english, then (generally speaking) you're at a disadvantage because you can only communicate with a subset of your population.
I'm not saying there's a problem with deaf people, any more than there's a problem with anyone else who simply doesn't happen to know the languages of some of the people around them.
Your comment not only perpetuates this totally false narrative that there’s a “problem” with ASL but it makes it sound like Deaf people have chosen only to socialize among themselves when the reality is that we have built a world that makes communication difficult for Deaf people. It doesn’t have to be that way: https://icyseas.org/2014/01/12/marthas-vineyard-deaf-people-...
You might find the Deaf mythology(?) of Eyeth interesting: https://www.nytimes.com/2021/10/10/opinion/deaf-population-i...
I don't see any "problem" with deaf people. I see ASL as, effectively, a foreign language; and it makes sense to me that in general, you're able to live more effectively when you can speak the same language as the people around you.
When I traveled overseas to a spanish-speaking country, I learned spanish so that I could communicate with the people there. It would be unreasonable for me to show up as the cliche american tourist and expect the spanish-speaking people there to learn english so that I could communicate with them.
Edit: got annoyed.
edited: the below sections refers to something that has been edited out
I understand that it might feel harsh, but to me, it looks like you're just being corrected and rebutted just like you said you'd be open to.
Rebuttals aren't just one-and-done. You have to be able to actually defend your position, and that requires a lot more than just one statement.
lmao it is absolutely you that's what we're trying to tell you
ASL isn’t a “foreign” language. It’s an American language. In fact, it’s more “American” than English. People who communicate via ASL aren’t foreigners on holiday. They’re our friends and family and neighbors. You are absolutely correct in your assumption that being able to communicate with those around you is important!
Imagine if everyone around you just refused to engage with you verbally and would only communicate via text messages. If you’re speaking they completely ignore you until you write it down. If they’re speaking while you’re around it’s always in whispers so that you can really only get bits and pieces. How would you feel? Annoyed? Excluded? Like you’re not able to fully understand conversations?
Your travel analogy doesn’t hold water because (in addition to Deaf people not being foreigners or guests!) you seem unwilling or unable to understand that (1) hearing is a critical component of effective verbal communication and (2) Deaf people can’t hear. Again, I’ll chalk it up to ignorance rather than ill intent but your analogy as a whole is pretty gross to paint Deaf people as entitled, unreasonable, and demanding. Your analogy is akin to suggesting it’s unreasonable for me, a person who can walk, to rollerskate everywhere I go and it would be unreasonable for me to expect curb cuts just so I could rollerskate everywhere so therefore it’s unreasonable for people who use a wheelchair to expect curb cuts. If someone who uses a wheelchair wants to use the sidewalk they should just try fucking walking right? I don’t know why no one has thought of that!
> ASL isn’t a “foreign” language. It’s an American language.
Elsewhere in this thread you mentioned that "ASL is a complete and distinct natural language in its own right," and that's the sense I was trying to convey by calling it a "foreign" language. It's as hard to acquire as a foreign language, and carries the same benefits. And it fit the analogy.
I appreciate both of your analogies, but I think they aren't quite fair. Not going to the massive effort of learning a new language isn't really the same as 'refusing to engage'. And the analogies seem to understate the effort involved by everyone who is not deaf to learn ASL. It's just not as easy as installing curb cuts (I could probably do one of those in an afternoon...become conversational in ASL? not so much).
I guess, to me, it comes down to the level of effort being asked of me. I'm happy to be accommodating, up to a point. Let's engage in whatever medium is most suited. But asking me to put in the massive time and intellectual commitment of learning a new language is past that point. And I think that's the case for a lot of other people as well; maybe most.
> If someone who uses a wheelchair wants to use the sidewalk they should just try fucking walking right? I don’t know why no one has thought of that!
I had a relative who was deaf. An engineer. He learned to lip read, and that gave him more freedom to work with others and be effective in communities that did not know ASL.
I tried to explain the cruel history behind your callous remark.
I tried to give you an opportunity to empathize with how a Deaf might feel.
I tried to offer some perspective as to why the views you’ve expressed are narrow-minded.
I tried to educate you.
But you’re just here to argue. You don’t want to put in the effort so it’s the Deaf community’s fault they can’t effectively communicate with you and people like you? Spanish is easy enough to learn for your trip but ASL is too difficult to learn? You know one person who lip reads so everyone must be able to?
You’re not even willing to consider for a moment that your gut-opinion or the second-hand experience of the ONE person you know might not be universal. I don’t have anything else to say to you except that I hope you take a few minutes to read about ASL and Deaf culture and history and take a few minutes to reflect.
Or you're just here to educate, and if someone questions you, you scold them for moral inferiority.
Yet you seem to think everyone is obligated to learn sign language...
>Spanish is easy enough to learn for your trip but ASL is too difficult to learn?
Spanish is NOT "easy enough to learn"; perhaps some basic phrases for tourists, but certainly not fluency. It takes years to master a new language for most people, if ever. Spanish is relatively easier for English speakers than many other languages, such as Japanese, but it's not something you're going to become conversational in quickly unless you're gifted at language learning.
The idea that people are somehow "cruel" and "narrow-minded" for not learning a particular language that isn't the dominant language in their region is the most ridiculous thing I've read all day, and that includes the comments from North Korea supporters here.
A system of writing can be used for any language. (Have you ever wondered why English, a Germanic language, or Vietnamese, an Austroasiatic language, use a writing system based on the Latin alphabet?) Any language that doesn’t use a system of writing can use a system of writing. It’s not like it’s some sort of existential impossibility.
If these can let people hang out and participate without having to actively track each speaker in a group setting it will go a long way.
It would be a technological Babelfish.
The form factor of Google's AR glasses look much much closer to a normal pair of glasses than the glasses in the top video (which look like heavy sunglasses with a wire connecting to your phone)
You posted a CNET link, here's a direct link to Google's: https://www.youtube.com/watch?v=lj0bFX9HXeE
In the description it says: "This device has not been authorized as required by the rules of the Federal Communications Commission. This device is not, and may not be, offered for sale or lease, or sold or leased, until authorization is obtained."
So I guess they might be waiting for that?
Or, if it had OCR capabilities, you could just hold a sheet of paper in front of yourself and say "what's this?" and it would explain the text to you.
As just some guy on the internet, can I buy one of these and write a hello world to have text show up in front of my eyes of my own choosing? Does it have an API or will it?
I do wish there had been a real-life shot showing how the text appears to the wearer, though.
I have an iPhone. How does it work with a device as part of the deal?
2. If these are the ones I have heard about before, all the speech to text is done on the device not in the cloud. This is for privacy reasons. Means it needs a bit more bulk for the gear to work.
I'm thinking a tablet that does this would be ideal for in-home use by some elderly relatives. AR glasses won't be suitable for some time for the elderly, as it seems eyesight starts going before hearing.
Travelers, translate Spanish to English etc.
Reconstruct voice in noisy environments.
???