Be My Eyes’ AI assistant starts rolling out
bemyeyes.com
bemyeyes.com
IMO, this announcement is far less significant than people make it out to be. The feature has been available as a private beta for a good few months, and as a public beta (with a waitlist) for the last few weeks. Most of the blind people I know (including myself) already have access and are pretty familiar with it by now.
I don't think this will replace human volunteers for now, but it's definitely a tool that can augment them. I've used the volunteer side of Be My AI quite a few times, but I only resort to that solution when I have no other option. Bothering a random human multiple times a day with my problems really doesn't feel like something I want to do. There are situations when you either don't need 100% certainty or know roughly what to expect and can detect hallucinations yourself. For example, when you have a few boxes that look exactly the same and you know exactly what they contain but not which box is which, Be My AI is a good solution. If it answers your question, that's great, if it hallucinates, you know that your box can only be one of a few things, so you'll probably catch that. Another interesting use case is random pictures shared to a group or Slack channel, it's good enough to let you distinguish between funny memes and screenshots of important announcements that merit further human attention, and perhaps a request for alt text.
This isn't a perfect tool for sure, but it's definitely pretty helpful if you know how to use it right. All these anti-AI sentiments are really unwarranted in this case IMO.
I've written more here https://dragonscave.space/@miki/111018682169530098
This won't entirely replace human volunteers, but these models get rapidly better over time. What you are seeing today is a mere toy compared to the multimodals you'll get in the future.
Currently there's no model trained on videos, due to large size of videos, but in the future there will be video-capable models, which means they can understand and interpret motion and physics. Put that in a smart glass, and it can act as live-eyes to navigate a busy street. Granted this will take years to bring the costs down to make that viable.
Videos may not suffice. Videos are 2d, with 3d aspects being inferred from that 2d data, which is an issue for autonomous driving based on cameras. A proper model for AI training would be 3d scans rather than videos. The best data set would be a combination of video and 3d scanning. Self-driving cars which might combine video with radar/laser scanning may one day provide such a data set.
There is talk of a 3d version of Google streetview, one using a pair of cameras to allow true VR viewing. That might also be good training data as it will capture, in 3d, may street scenes as they unfold.
But in a higher-acceptable-latency, lower-risk environment like this, I am actually quite bullish on camera alone methods. Video understanding has come a long way.
We had enough drama with BeMyAI refusing to recognize faces (including faces of famous people) as it were. If sighted people have the right of accessing porn and sexting, why shouldn't we? Who should dictate what content is "appropriate", and what about cultures with different opinions on the subject?
OpenAI should dictate that, because GPT 4 belongs to them. So they decide what kind of service they're interested in offering.
There will be plenty of other powerful LLMs that can be used in the near future. Some will be more restrictive, some will be less. If you want fewer restrictions, you will be able to pick one that offers that for you.
Extremely optimistic take. What tends to happen is you get centralisation, and regulatory capture ensures the largest players dictate what is an acceptable to the incumbents who wish to do things differently.
I mean in theory you can go set up your own social network or video sharing site with whatever rules you like, but you should assume government regulators and big tech will attack you if you do so and believe in the principles of free speech, or simply wish to create a safe-space for conspiracy theorists.
Please don't feel like you're bothering us. I've had this app for years and absolutely cherished the few calls I've gotten. I get really bummed if I miss a call.
The people who sign up to help (such as myself) want to help. Honestly, my frustration is that I don’t get asked enough.
That can happen to anyone. Some blind people are introverts and don't want to talk to random strangers all the time.
Also, while the vast majority of volunteers have the best intentions and try hard to be helpful, you never know what you're going to get. Some are way too chatty, some offer unsolicited advice.
Calling a human being is much higher-friction than just opening an app, if that weren't the case, we' all still be calling restaurants instead of ordering on Uber Eats.
I also would prefer not to call volunteers at night. This isn't much of a problem if you live in the US, as the app is segregated by language, not country, so you'll probably find somebody down under. I have an additional complication of having to deal with foreign language content because of where I live, so English often isn't good enough for me.
Just wanted to say, not talking whatever you feel you should do, but me and others I know don't feel at all like you are 'Bothering a random human multiple times a day'. I personally feel extremely lucky when I get a 'call' and am able to help.
Not much else regarding your comment, just wanted to let you know that I have had very satisfying 'calls' and I am always looking for the next one and always happy when I do.
A few companies are trying it. Last weekend, I tried one, and it was not yet polished, but good enough to start. It recognized me (my friend already has me on his phone contact).
I’m trying to help a friend with his non-profit initiative, bringing the cost to about a third or even a-fourth of Google Glass[1]. After reading the article and the other First Impressions with GPT-4V(ision)[2], it is apparent that this can be possible much simpler and soon enough. The rumor about Ive[3] and Sam[4] talking about AI in Hardware has already given me some good hope.
If anyone else does AI-enabled hardware assistance for blind people and has devices that will be less than $500 a piece in retail, I’d love to talk and introduce them to the right people.
1. https://en.wikipedia.org/wiki/Google_Glass
2. https://news.ycombinator.com/item?id=37673409
I'm the Founder & CTO of Envision. We're building EXACTLY the product you're describing. Envision Glasses is a bunch of computer vision tools built on top of the Google Glass Enterprise Edition 2. We've more than 2000 visually impaired people, across the world, using the Glasses in more than 30 different langauges.
You can check out more informaion here: https://www.letsenvision.com/glasses
P.S: I know the EE2 has been discontinued but we've been working closely with Google to ensure current and future demand is met. We're also experimenting a lot with other exciting off-the-shelf glasses that I can't talk about here but I'm super excited for this whole glasses + AI space!
It was one of the best apps out there for (instant) text recognition, and I was pretty happy to pay for it, but since it went free, it's really not the same. There's nothing else like the old Envision out there.
Also, not an issue for me personally, but dropping support for Cyrillic in the middle of a brutal war in Ukraine, in an automatic update released with no prior warning, was an asshole move if I ever saw one.
I think without a major consumer product as a platform, it will never reach the number of people it wants to help. AS in example take what the iPhone did to the market of braille note takers.
It's still not glass integrated - but the phone version has been out for a while ( https://www.microsoft.com/en-us/ai/seeing-ai ) - previously discussed when it was released https://news.ycombinator.com/item?id=14774167 (6 years ago, 133 comments)
* It can help reading a menu. But not just straight linear reading... I told it I prefer veggie today, but would also eat something light with meat. So it highlighted the veggie options for me, and in one case even tried to guess the type of dish from a photo which lacked a text description. * I had it search for cobwebs on a ceiling. Vacuumed them away, and took a second picture, asking if they were gone now. * It told me a houseplant of mine has yellow leaves, and told me which by describing a path via the stems.
In general, the scene description is very good, and the built-in (implicit) OCR is just a game changer. It is like having a real human reading something to you. Typically, you dont want the human to read all the text on the page, typically you are only interested in some detail, and usually instruct the helper about that so they dont have to read everything to you. The same now wors with an AI, with consistent quality. Volunteers are great, dont get me wrong, but this comes with a lot of problems. You sometimes have to request help several times, because frankly, sometimes you simply get people who can not help you due to their own abilities. Also, the video feed is demanding on connectivity. If I try to read a menu in the train with volunteers, most of the time this will fail due to "blurry vision" and the camera dropping out. Sending a single picture every few minutes is totally OK in these situations.
Of course, they explicitly push responsibility to the user by saying things like:
> No. Do not use Be My AI for scanning medicines, reading dosages, or other safety issues
God forbid that we acknowledge this tech is fundamentally flawed. There’s money to be made!
In some sense, I was a non-AI solution that enabled them to do their activities of daily living.
I'm pretty sure they'd have preferred not to be dependent on me, though. I imagine that's true here, too--no more need to edit your behaviour/filter your thinking/navigate interpersonal stuff just to do something minor.
Humans have well understood failure modes, and society has an existing framework for handling issues that arise between humans.
I wonder who will be held responsible when Be My Eyes AI tells someone to step into traffic? My guess is no one.
> Be My AI is perfect for all those circumstances when you want a quick solution or you don’t feel like talking to another person to get visual assistance.
This summer I used Google Translate to pick medicine in Italy and it was pretty good at translating labels - definitely better than pharmacist who did not speak English at all.
By the way, lots of people die in US because wrong medicine was dispensed - and that has nothing to do with AI. People are imperfect and many drug names are long, incomprehensible and easy to confuse with each other.
Lots of blind people have their homes set up with everything carefully arranged and no hazards. They prepare food, clean their house, commute to school, to jobs, and to appointments via familiar routes with no dangerous intersections, or they might use buses, trains, and Uber / Lyft just like millions of other people who aren't blind.
Being blind can be full of annoyances and frustrations, but I think it's a stretch to say that it's "dangerous".
That’s wrong and immoral. The world is full of risk. As individuals we choose to accept some risks and move on with our lives. I think it’s appropriate that blind people are given the same opportunity
I don't believe that Be My Eyes is a for-profit entity.
AI right now ain't no god, but the gap between AI mistakes and human mistakes isn't that wide.
This argument that "humans are flawed too" is not useful, we know humans are flawed, that's why we want tech to fill those gaps, not flawed / potentially dangerous tech, actual tech like how a plane flies you to your destination 99.9% of the time.
1. Take a photo of a piece of tech: AI can not only describe what's going on but tell you exactly how to fix it (e.g. "it's showing a fatal error. hold down the power button for 10 seconds to reboot it") 2. Take a photo of clothing: AI uses consistent, neutral language to describe garments, while humans would all describe the same item differently based on where they live, their own personal style, their region, etc. 3. Take a photo in a major city or near a famous landmark, AI will recognize it and tell you about your location, whereas a random human has probably never been there before 4. Take a photo of an insurance bill, AI can explain what some of the technical terms mean if you don't understand them 5. Take a photo of a menu, ask it to summarize all of the vegetarian entrees and it will list them in seconds ("one veggie burger, one stir fry with tofu, three sandwiches and all of the salads"); a human might take a minute or two to carefully read the menu and give a much longer answer
To be fair, humans make mistakes at an incredible rate. Even the ones who make life or death decisions.
I'm talking about the constant stating of the obvious, if these models weren't useful we wouldn't deploy them. It's like the whole internet constantly making the statement "cars are better at covering distances in shorter time frames than humans"...we know, that's why we build them and pay a bunch of money for them.
On the plus side, I suppose in this application with the humans there's always a risk of a malicious volunteer outright lying, which presumably isn't a risk with the AI solution.
I've been playing with Be My AI, it's excellent. It's got a really good prompt that enables it to be really descriptive and helpful with very little hallucination. It's actually a lot better at describing most photos than the average person is.
Making the world more accessible to blind people is a big challenge. There are no perfect solutions, but the solutions we do have today are enabling millions of blind people to have full careers and live independently.
This is a great improvement. It's not perfect, but neither were the previous technologies that blind people have been using. Don't let the perfect be the enemy of the good.
I suspect most of the anger in this thread is just the usual anti-AI stance, but it honestly feels extremely patronizing in this context.
I don't know about other countries, but here in Switzerland I've had the app for five years and got only 4 requests for help during that time, which made me think that there were way more helpers than help-seekers. But I suppose they wouldn't be adding AI if that was the case.
Yes, about 15:1.
https://www.bemyeyes.com/blog/4-million-volunteers-strong
https://www.bemyeyes.com/blog/what-would-you-do-with-5-milli...
> Be My AI is perfect for all those circumstances when you want a quick solution or you don’t feel like talking to another person to get visual assistance.
It presumably will be getting used in low risk situations.
Please dont project the typical AI-angst into this assistive technology. It doesnt deserve this, especially from people without hands-on experience.
Two things I've helped with:
Wrapping a gift by helping them orientate the gift correctly (I don't even do this!)
Preparing a meal by reading ingredients / labels.
On the other hand, I’d like to be able to test its ability to describe graphics to me. If it’s able to turn graphics into an accessible table, I can browse with my screen reader that would be revolutionary.
Original comment: Yeah, this wasn't always the case, until recently. People were using it to describe their kids, spouses, etc. Pissed a lot of folks off when they disabled it. I never even thought of using it for that, so now that I realize I could have, it kinda' pisses me off as well. Honestly, I never cared too much what they looked like, but it would have been interesting to hear the ChatGPT viewpoint. :)
But definitely not something I could have asked a fellow human about without it seeming really weird, and being confident I wouldn't get a biased answer. Though it's probably unrealistic to expect that the AI wouldn't also give a biased answer.
As we quickly learned, the operator would say anything. Anything. So for a few days we would call each other with the operator and never were able to find the limit. I have great respect now, those operators were inhumanly stone faced, and respect that the system was perfectly transparent. Nothing typed was hidden or obfuscated.
Are there a large number of users that feel like they are wasting volunteer time with menial tasks?
As for menial tasks, I could definitely see people wanting to use this instead of calling a stranger for more personal matters -- at least initially.
Not to mention I imagine the carbon cost of pushing this through AI is much higher than just using humans.
This sort of dismissive comment is very anti-social and full of hidden hatred. Projecting your squarrels with a company onto people that really need the help provided.
And before you click, I am that pissed because I am blind. You have absolutenly no idea what that means, and what BeMyEyes and BeMyAI did for us. Just go home and hate someone else please.
That's the only death of a self-driving car. That's the only one. There was a safety driver with their hands on the wheel, and they didn't see the pedestrian either. And that was an Uber self-driving car, and they've since canceled that project.
While incredibly tragic, it's not at all fair to say that self-driving cars are constantly plowing into people.
Waymo has driven more than 1 million miles with no human injuries. That's dramatically better than any human driver.
Cruise is a close second in number of miles, also with no human injuries.
Perhaps critical thinking is unrelated to sight?
I didn't read any hate towards anyone in the previous comment?
There are a lot of things that are just tech demos. This is far more than that.
Considering how often I’ve seen the complaint from your parent post, it’s quite clear people don’t mind. Quite the opposite, they’d embrace the opportunity. Maybe the people who need assistance don’t realise that, but again, that complaint is quite common. I’d like to help but never signed up specifically because of that surplus.
So they had a solution based on humans who are eager to help and are replacing it with an automated system which when mistaken can have disastrous results and cause personal injury. Seems odd to me. A humanised approach is often seen as a positive and this cuts it out without necessity.
All that said, I don’t have any insider information. Perhaps the people who need assistant do prefer talking to a machine.
When I received my first (and only) BeMyEyes request I spent the first minute or so figuring out how to work the app and the video delay.
“Let me turn the volume up” “Oh wait it’s on my AirPods” “Could you move it more to the left?” “No your left.” “No; not that far”
I’m quite confident that I was a less than optimal assistant and an AI might’ve well done better than me.
Humans are slow and unreliable. AI is fast and consistent.
Both are imperfect. You shouldn't rely on either one for something life-or-death. Nobody is using them to decide if it's safe to cross a street.
Thanks for being a volunteer. But please dont judge blind users if you apparently can not put yourself in ther shoes...
Yes, one reason is that some blind users dont want to waste volunteer time. Another reason is that volunteers are different, but the performance of the image AI is predictable. Another reason is that the AI OCR is fast, can also translate, and, surprise, the text is easier to handle later, for copy&paste.
Besides, the performance of the AI describing pictures is, sorry to say, a little bit above what the typical human is willing or capable of doing. IOW, some humans performs worse then the AI. Also, camera access is different from picture taking. I use volunteers when I need more interactive help, but I totally prefer the AI when I just need a single pic.
I think we can imagine an AI could describe the screen, and even find non-language visual elements if asked explicitly, like arrows or turtle vs hare icons etc. But is it ready to have shared context of how people need to interact with that UI?
And what if it is not? Does the AI have to be useful for every single possible use case before it can be used?
You are reacting as if they are proposing to remove the already existing venues of help. When I see no sign of that.
I know blind people who are making just as many human Be My Eyes calls as they were before, but they're using Be My AI even more, for things where they wouldn't have even bothered to use the service at all before - because the AI is so fast and convenient.
What we found is that a small cohort of our users were in some way visually impaired and were very reliant on the app. We ended up consciously deciding to focus exclusively on this cohort which is not a typical product management strategy (assuming your trying to max $) but the remaining enhancements we did made some lives better and we were happy with that.
Adding AI for basic things would probably be a game-changer for blind people who have lots of questions, or embarrassing questions, or want to read their credit card number or something.
If you had to choose a pathway to fix the blindness issue, which one would you choose? Why?
See: https://www.technologyreview.com/2022/08/11/1057576/bioengin...
But we don't need to choose. We are not living in some Age of Empires like game where the town centre can only develop one item from the tech tree at a time.
The set of people who can develop this are entirely different from the set of people who can work on a bioengineered cornea. The pot of money this is financed from is not the same as the one we would be financing those other projects.
Finally there are many separate biological problems which can cause blindness. A bioengineered cornea might help with some of them, but not others. Even if we would have a truly amazing and cheap and reliable bioengineered cornea we would still have blind people whom it would be unable to help. Because of this it is worthwhile to work on both.
This is about understanding why your choice. If your answer is, I'd give one pot to each, and one of them is for the short run, that is okay.
But I'm really trying to invite answers and understand the downsides and upsides of each.
Pros: 1. AI is way more proven of an investment. There's an extremely clear path forward where money = better vision capabilities.
2. AI is extremely cost efficient and cheap to mass deploy. Blind users can get access with a cheap monthly subscription of $20, easy to afford for vast majority of developed world. Compare this with a cyber-cornea, which even if it worked, even if it didn't cost $10k to make one each, would still have massive costs just for installation, and benefit a tiny amount of relatively wealthy disabled.
3. AI works for universal blindness. A cornea can only fix problems inside the eyeball itself.
Bioengineered corneas may restore sight to people with keratoconus.
AI description tools can provide assistance to people with impaired vision.
One is an experimental medical procedure, one is an app. One is a fix, one is a tool. One is specific, one is general.
What could we learn from such a comparison?
As a volunteer, however, it makes me a little sad. I enjoy being helpful and valuable.
I'm also curious about purely AI based versus LiDAR solutions.
Also, note that many blind users access Be My Eyes using a refreshable braille display.
“You are standing in an open field west of a white house, with a boarded front door. There is a small mailbox here.”
It would be relatively lightweight but still "realistic"
I'm not well versed on Ascii to image conversion, but I guess you could use AI to segment and recognize objects too complex to recognize easily and adjust "image" generation accordingly
Now DALLE-3 is good enough, problem is if it'll be cheap enough. 100% there are revolutionary new forms of adventure games being developed based on it.
[1] https://dzlab.github.io/notebooks/flax/vision/diffusion/2023...
[2] https://www.reddit.com/r/LocalLLaMA/comments/11wwwjq/graphic...
[3] https://huggingface.co/spaces/jbilcke-hf/VideoQuest/tree/mai...
Not yet, OpenAI is too preachy , if your adventure would bring you into a bar and try to drink something you will get a few paragraphs about "alcohol is bad", and is also very "child limited/targeted" so if a monster spawn , the hero will most likely befriend the monster will love and they will live happily ever after.
The open uncensored ones are still WIP last tiem I checked, they either forget the conversation from a few moments ago, or the training made them dumb.
If I am wrong then someone tell me where I can try a demo for such a text adventure that has memory and is not crippled like it targets "small american children".
I normally test LLMs by providing a list of bizarre "weapons" and asking how I could use them to defeat an improbable beast.
It turns out that an enormous dust mite the size of a car needs a disclaimer when your weapons are a comb, a plastic horseshoe, an etch-a-sketch, a baseball glove, and a worn copy of the farmer's almanac. That being said, ChatGPT still tends to wins the creativity test.