Humane AI Pin
hu.ma.ne
hu.ma.ne
The Humane AI Pin Launches Its Campaign to Replace Phones - https://news.ycombinator.com/item?id=38207656 - Nov 2023 (130 comments)
I think this is easy to dismiss at first glance, but I genuinely believe they're trying to think about a new mode of interaction. The idea that "the computer will disappear" is probably accurate in the long term. Except for content delivery (reading, photos, movies), most tasks we achieve via computers and phones do not strictly require a screen. It's probably a good thing if computers did a better job of getting out of the way, and stop so loudly disrupting human interactions.
Whether this will be the solution is unclear; the privacy/creepiness angle is still real with an outwards-facing camera. Latency and battery life limitations might be too significant. The cost will be a non-starter for many (it is for me).
But I'm still impressed because there was a vision here. The conversational interface has never worked before for many reasons, but that does not mean it cannot work in principle, or that the ideal implementation would not be spellbinding. I'm glad they're trying. Also, the laser display is neat!
X (doubt). There are unfortunately only 5 senses that our brains can interact with the outside world, and visual ways are the most information dense and the easiest to utilize. The screen isn't going away anytime soon.
Projector to me are same as screen - they've been around for as long too.
Though I do look forward to direct computer-brain interface, like introducing a 6th sense.
I too would be interested to see an enumerated list of over 100 senses.
"Sight" split into rods for brightness sensitivity, and cones, each of which is deicated to one out of red, green, and blue. green is wider gamut of color than the others because there is a lot of green in nature. These sensors are fully independant of each other for the most part, although there is minor overlap between cones which is what we call other colors (yellow etc)
"Taste" Again split into different specialised papillae sensors. I dont remember so well, but its something like foliate for sour sensing, fungiform for salty, and vallate for bitter/poison. There is also sweet I dont remember the name, and some argue for umami
"Touch" There are an ungodly number of very distinct senses that go into touch. From more abstract ones like pain, heat/cold, moisture (not evenly distributed around body, for example have to touch things to lips to distinguish cold from wet), proprioception for joints (arguably an independant sense for each joint, or at least each "kind" of joint, because the biological mechanism is different for ball joints to saddle joints etc as well as specialised proprioception for eyeballs, tongue etc)
Then in actual touch touch there is Ruffini corpuscles sensing skin stretching and slippage of objects past the skin
Merkel discs, which senses pressure applied to the skin and low frequency vibration
Meissner's corpuscles, which sense vibrations in middle range. They are very sensitive and allow very slight sensing of tiny impulses such as picking up an insect's wing
Pacinian corpuscle sense extremely fast vibration which among other things allow the distinction between "rough" and "smooth" surfaces (by mechanical movement causing vibration)
There are also free nerve endings sensing stuff like itching and bruising.
Hair foillicles also sense movement and stretching of the hair they are attached too, which provides more touch data. Incidentally this mechanism is also used for balance and hearing via really complicated interactions of tiny hairs in the ear.
"Smell" Smell is fiendishly complex, it actually is more akin to the way antibodies in the body are made in the sense it consists of thousands (and millions) of specialised sensors made to "fit" and attach to individual compounds, so there are almost limitless individual senses of smell
There is also a whole lot of internal sensor data for things like breathing (you know when you are short of breath), digestion you know when you are full, or when you are craving one of a number of things sweet salty etc), bladder control.
This is mostly off the top of my head and i'm certain i'm misremembering some of the subtlties and a whole bunch more senses both obscure and immediately recognisable ones to any owner of a human body
On the other hand, in the context of the discussion, it's hard to support the argument that you can count each colour channel separately just because the biological mechanics differ. You can't actually triple the amount of human-perceptable information by going from a monochrome to full colour display.
The point remains that we've plucked the low-hanging fruit when it comes to high-bandwidth human senses (or meta-senses if you insist on being pedantic). No one will buy a PUI (pain user interface).
You absolutely LOSE perceptible information when you lose one of then channels, like in color blindness.
Other types of interfaces do exist, for example ive worked with vibration motor arrays placed on the skin for various purposes such as assisting in guiding the arm of a patient to target a specific point (vibrate on side closest to target) etc. We also worked with pads of electrical patches that pass small currents through the skin to produce a distinct sensation, like pain but barely at the threshold of being noticable. These were used for first responders, placed along the side of the torso underneath the clothes with flat profile, allowing them to have handsfree silent communication with low bandwidth. Something like "up up left left" being pre-agreed to mean leave the structure now etc. Another fun one I wanted to mention is in-mouth joysticks controlled with the tongue for quadreplegic patients to allow them to move a wheelchair or robot arm to regain some small independance (might seem like it would be uselessly hard to achieve anything with an arm controlled that way but the emotional impact of independance can't be understated for such people, even a simple task can be very meaningful)
They won't be as good as screens or audio unfortunately. But they can exist. Even braille screens and keyboards exist as a nice product and are reasonably high bandwidth.
True, part of what makes them cool is that your proprioception more or less agrees with the virtual hand that you see in your headset, but that's just window dressing. The computer has no way to control that.
Even on the vision front, we have rods and cones that works differently to generate ONE vision.
This is entirely semantics.
The next generation of devices that incorporate some of these features might be more successful.
I don't think you're wrong, but it's funny that we aren't as concerned about everyone walking around with outwards-facing phone cameras.
I myself never felt like taping my camera, I feel like if someone pwned my system I would be much more worried about the leaked audio.
I've always said that privacy is an illusion, the usual example I give is: "You're lying in bed with the curtains drawn, you see a shadow fall across the curtains that looks like a person standing outside. Do you, or do you not have privacy?"
If the shadow turns out to be a person peeking through the curtains, then you don't. If the shadow turns out to be primal brain + tree shadow then you do. Schrodinger style.
Privacy is probably best described (as it sometimes is) as a "sense" of privacy I guess.
> The conversational interface has never worked before for many reasons, but that does not mean it cannot work in principle, …. I'm glad they're trying. Also, the laser display is neat!
So I did a lot of work over the years to research voice UI/UX and I’m very skeptical about this, even with the LLM stuff. I think an LLM was missing from the Siri/alexa era to transform it from “audio cli” to “chat interface” but there’s a few reasons besides that it didn’t catch on.
The information density and linearity of chat, voice especially, is a big problem.
When you look at a screen, your eyes can move in 2 dimensions. You can have sidebars, you can have text fields organized in paragraphs and buttons and bars etc. Not so with chatting - when you add linearity (you can only listen to or read one thing at a time, conversation can only present one list at a time) it becomes really slow to navigate any sort of decision or menu trees. Mobile-first have simplified this of course, but it’s not enough. Reading TTS becomes even slower to find the info you care about. It’s found a place for simple controls (smarthome, media, timers, etc) and simple information retrieval (weather, announce doorbell, read last text). Then there’s the obvious problem of talking out loud in public, false response recognition etc which are necessary evils of a voice UI.
I think the best hope for a voice device like this is to (as they’ve done) focus on simple experiences like “what’s I miss recently” and hope an AI can do a good enough job.
The laser display might help with presenting a full menu at once (media controls being an easy example), but it probably will end up being a pain to use (eg like a worse smartwatch).
Honestly though, my biggest hesitation (which could end up great) is the “pin” design. It’s novel, especially with the projector, but how heavy is it and how will that impact the comfort of my clothes? What about when wearing a jacket or scarf? Will this flop around while walking? Etc.
If a science fiction author was writing it, the need for stiffer fabrics to support chest cameras would synergize with a neo-Victorianism in generation alpha. (Formal button-up shirts and higher necklines for enforced modesty)
But yeah I've been thinking that too. "Oh, put my coat on - better spend 30 seconds messing around with my pin" [...] "Ahhh back in the office. There goes another thirty seconds moving the pin so it can film me looking at a screen for four hours"
And yeah, I feel like the weight would definitely pull my jumper or t-shirt out of shape, and make things like my collar/neckline look out of whack. Maybe they'll bring out a range of clothes suitable for it, or suggest you wear a coat indoors like the woman in the video is doing.
Let's not forget the value of non-linear input. Good search terms are often constructed rather than spilled forth. Sometimes I enter search terms, read it and realize that it's like to return unrelated results and need to modify it. By the time I realize this while speaking to an AI it's already spitting out the wrong information.
This leads to a need for altered interfaces that allow these scenarios to be accomodated. This is v1.0. Let's see where it goes.
Even now - clicking through some insurance company's website hierarchy to find something out is insanely painful.
But even for researching things that we should probably care about enough to do it ourselves, correlating different sources of information or working through abstract/ambiguous problems... the vast majority of ordinary people will 100% take the easy way out and let LLMs do most of the thinking for them. Even with free GPT-3, people are unflinchingly having LLMs solve problems they don't want to think about too deeply. What they pay for, with occasional inaccuracy, is more than offset by convenience.
Maybe, but I don’t know if that day is here yet. I think “most people” do actually consume information. Like reading an insurance company’s website is pretty rare compared to things like using the Amazon App. Like it’d be hard to consume a list of 5+ push notifications via voice if you had to listen to them 1 by 1 instead of skimming them in a list next to their icons.
Even simple things like scrolling through a list of songs becomes painful. I have like 10k songs in my (streaming) library Sometimes I randomly scroll through it to find old music. That sounds impossible on voice. I’d be stuck with “shuffle” mode.
Being able to summarize and search text conversations via voice queries from their demo would be nice, but today that’s a task that you need a screen for.
The demo video shows the man buying a book online via voice after holding it up to the camera. How often is that the online shopping experience? I can’t imagine shopping without a screen 95% of the time.
we may not need it but we certainly prefer it. People went completely voluntary from voice calling to texting and within texting to ever terser forms to the point were an entire website was built around a short character limit.
Except for people with disability I have not really seen a single case where that tendency towards compactness is reversed in communication.
Why though? Computer requires attention, which pretty much rules out doing something else while using it, except perhaps when passively listening to a podcast (which doesnt really qualify as computer use). Even though we may see new mediums, the mode of interaction will remain similar to that of a book
And that is not this. Talking out loud every few moments with verbal commands do a device is way more annoying that someone looking at and typing on a phone
That said, I agree with you at a glance it's neat. I think in reality though it's a poor idea given how often people need to give a verbal command.
Also bullish on hand gesture control. Maybe most stuff will eventually become jutsu level fancy hand movements lol. What a time to be alive. It is easy to remain grateful in this age of rapid progress.
If you want the computer to disappear, why not a better smartwatch? Or glasses, this time without the sci-fi gadget look? Both could support the exact same featureset but with a screen.
You'd think your tech demo would check to see if your AI was hallucinating!
You’d think they’d have learned their lesson after Google Bard’s hallucinated demo!
Cant believe they left this stuff in.
Gotta at least make it seem good in the commercial, this ended up being the opposite of a sizzle reel
I’m not an AI detractor. I use it and really like it. I just don’t like it for information like this. Anything where the response needs to be verified yet is very brief makes no sense to bounce off of an AI, in my opinion.
Alexa was the closest to achieve significant usage since you can use it within the privacy of your home.
For voice UIs the non clear boundaries on what you think it can or cannot do is also a huge hurdle. After you get a couple “sorry I cannot do that” you stop using it
Combine that with areas like GPT Vision, (GPT?) Whisper, etc .. it'll start feeling a lot more natural here very soon i suspect.
TBH i'm surprised Apple isn't pushing this much harder. They tout Siri so hard but it's just worthless to me. It feels like apple could make a AI Pin like this, but visibly from the public side i have zero idea that the're even working in this space. It feels like they purposefully watched the boat sail away.
edit: Sidenote, Pin + Airpods would be a nice way to interface more quietly too.
These next gen AI voice assistants are still a solid improvement over Google's current offerings, but they'll feel like a massive jump into the future for folks that have been stuck in Apple's ecosystem, and that's probably where the biggest opportunity lies.
I went through this exercise with GPT voice. It's an awesome capability, but other than perhaps walking outside, or sitting in my office, there's no other space where it feels "ok" to just spontaneously talk to something.
A grey area is when you perhaps have headphones in / on and it looks like you're in a phone conversation with somebody, then it kinda feels ok, but generally you're not going to take a phone conversation in a public area without distancing yourself from others.
There's a reason most casual communication these days is text rather than voice or video calls.
"Siri, lights to HALF."
"Siri, lights to HAAAAALF."
"Siri, LIGHTS TO FIFTY PERCENT!"
Silly how OpenAI could blow all voice assistants out of the water today, if they just added Android intents as function calls to the ChatGPT app. Yes, the "voice chat mode" is that good.
With alexa i can program if/then statements, like basically when i say X then do Y. If something like chatgpt requires the same thing then i don't see the advantage.
So LLMs today can do this a few ways. One they can write and execute code. You can ask for some complex math (eg calculate the tip for this bill), and the LLM can respond with a python program to execute that math, then the wrapping program can execute this and return the result. You can scale this up a bit, use your creativity at the possiblities (eg SQL queries, one-off UIs, etc).
You can also use an LLM to “craft a call to an API from <api library>”. Today, Alexa basically works by calling an API. You get a weather api, a timer api, etc and make them all conform to the Alexa standard. An LLM can one-up it by using any existing API unchanged, as long as there’s adequate documentation somewhere for the LLM.
An LLM won’t revolutionize Alexa type use cases, but it will give it a way to reach the “long tail” of APIs and data retrieval. LLMs are pretty novel for the “write custom code to solve this unique problem” use case.
> and to a lesser degree malformed output
What's cool, is that this isn't a huge issue. Most LLMs how have "grammar" controls, where the model doesn't select any character as the next one, it selects the highest-probability character that conforms to the grammar. This dramatically helps things like well-formed JSON (or XML or... ) output.
Yes, I was thinking about even something as if/then, which could be configured in the UI and manifest to GPT-4 as the usual function call stuff.
The advantage here would be twofold:
1. GPT-4 won't need you to talk a weird command language; it's quite good at understanding regular talk and turning it into structured data. It will have no problem understanding things like "oh flip the lights in the living room and run some music, idk, maybe some Beatles", followed by "nah, too bright, tone it down a little", and reliably converting them into data you could feed to your if/else logic.
2. ChatGPT (the app) has a voice recognition model that, unlike Google Assistant, Siri and Alexa, does not suck. It's the first model I've experienced that can convert my casual speech into text with 95%+ accuracy even with lots of ambient noise.
Those are the features ChatGPT app offers today. Right now, if they added a basic bidirectional Tasker integration (user-configurable "function calls" emitting structured data for Tasker, and ability for Tasker to add messages into chat), anyone could quickly DIY something 20x better than Google Assistant.
Of course this isn't fool proof, and there still needs to be work defining capabilities of systems, etc. (although these are tasks AI can assist with). But it's promising - "teaching" the system how to do new things is relatively simple, and effectively akin to describing capabilities rather than programming directly.
The basic idea is to instruct the llm to output some kind of signal in text (often a json blob) that describes what it should do, then have a normal program use that json to execute some function.
Even though everyone's seen AirPods by now, in those rare occasions when I'm on the phone in public, I feel compelled to have my phone out and vaguely talking at it, so it's clear I'm on a phone call and not a crazy person.
I'm curious if we would see similar usage with the pin, where voice commands in public are always performed with the hand up for the projection screen (it will still prompt looks, but hopefully be clear in context, "oh they're doing some tech thing").
Of course at this price point, it's highly dubious that we'll see anywhere near the ubiquitous market penetration of AirPods (which garner understandable complaints about the price point sub-$200, and that's with a clear value prop).
The other reason they are mostly impractical - keeping a charge. *wired* headsets were great in this regard, but then there's the wire, and now, there's the phone (that may not even support the wire?).
Still I would guess Meta Glasses or AirPods should be better to handle such whispered mode since microphones are so much closer. Would be interesting if Airpods had some contact mic that could pickup whispered sound inside your mouth.
Maybe the holly grail is to have something inside your mouth so you don't have to even make voice - device will figure out what you want to speak from how mouth and tongue movement - smart tooth braces anyone? :)
The issues isn’t “communicating with the device” it’s “communicating around other people”.
I have almost negative interest in having to recite the technical specifics of my web search to my phone on the train to work. I have even less interest in having to listen to the person next to me trying to do the same.
I don't see how what you are saying holds any water in that scenario.
It does not seem right to speak of a single reason. There are probably multiple. So, IMHO it would be more productive to come up with a list and put some weights on the options if you want to dissect this matter.
IMHO one very strong factor / important reason (one that you ignore) is the social context. Ie the reaction of others in the same physical space, as you start talking out loud, seemingly unmotivated.
Humans are social animals, and so the reaction of others to the actions you do tend to be very important to a large fraction of the population. What is acceptable in one context simply isn't in another. Also, the exact tolerances tend to differ with the local culture (here "local" is used in the sense "geographically/physically local")
It's not just about not annoying others here. In this case it's also about a thing as imprecise as "perceived self image". Some people (I'd argue, most people) dislike having the perception that others perceive them to be mentally unstable or rude. Most people need some kind of social acceptance for the actions they do.
One significant trait of some mental instabilities (as well as some drug induced behavioral changes) is that those affected will spontanously start talking in public. You will probably know the Tourettes Syndrome, and the alchoholic rambling about because these cases often imply quite rude and offensive verbiage and/or loud volume, but these are not the only cases.
People in general are well adept at detecting such anomalous behaviour as it is part of our insticts trained through Evolution. Also the uncomfortable feelings that observing this type of behaviour leads to will lead many to react with a "confront or escape" (aka. "fight or flee") response (a stress signal), which is not beneficial to social interaction in general.
TL;DR: If you speak out in public without a very clear and socially valid reason (speaking to an object is not that) you are not only rude to others, but you also cause them stress... and you will have to face the social stigma of being perceived as insane.
(edit: grammar/typos)
Except... this problem is known to be trivially solvable. After all, the very act of putting a flat rectangle to your ear makes talking out loud in public not just perfectly acceptable, but mundane and not worth paying attention to (subject to social norms dictating where it is or isn't OK to be on the phone).
As for talking to yourself signalling insanity... I'd hope that stupid and probably developmentally retarding idea died long ago, and the "talking to yourself out loud in public" subtrope being dead since wireless earphones got ubiquitous some two decades ago.
The modern reality is, hearing someone "talking to themselves" is normal, and 99.9% of times means they're on a call.
The point is it's not, though. As a society we have generally established that it is rude to be speaking out loud on the phone in public. Especially on the bus or the train or waiting for same or in the shop or at a movie or any number of other places. I genuinely think it would be easier instead to list the places where it would be okay (in a busy street, if you step to one side). Even in these places there is some expectation that you show a little shame to be doing it, as though you didnt want to but had to because the call is important
After watching the presentation, I am now curious about Humane’s thing though, but I’m still going to hold off for a bit because I want to see the failure modes first and I also don’t want to rush out and be one of the first to buy the brand new 3Com Audrey.
Talking to my cuff isn't going to make this better
If we could subvocalise with throat or other microphones/bone speaker then maaaybe, but I feel like it's better left to a brain interface and we should really just stick to touchscreen/typing interaction for now.
(Yes, Siri is not great today, but that will change very quickly with Apple working hard on their own LLMs.)
Cool project, but not something I imagine most people will want. Like Google Glass.
They even did the cringey stunt Google Glass tried and featured it on the runway during Fashion Week, as if that instantly makes something fashionable:
https://images.fastcompany.net/image/upload/w_1200,c_limit,q...
Though from the reviews I've seen (and as with so many Bluetooth devices), it's unusably terrible, and the battery only lasts a few hours.
A touch more seriously, the Narrative Clip:
https://en.wikipedia.org/wiki/Narrative_Clip
https://thenextweb.com/news/narratives-clip-2-wearable-camer...
This thing (the Humane AI Pin) is aiming to be a phone replacement, which seems like a really steep challenge given its limitations--how could it replace any of the things I use my phone for on the subway to work?
It’s a great point that if this modality becomes popular, then it should just be an accessory on top of iPhone or iWatch.
Think of the simple interaction of wanting to issue a voice command in public. Watch: Bring it close to your mouth, maybe cover both with the other hand to be even less audible to others. Humane: Smoosh your shirt up to your face?
(Also: I live in one of the sunniest places on earth — I simply don't trust that I'll be able to see light projections onto my hand when I'm outside.)
Anyway. All in favor of exploration and new ideas. Very willing to be proven wrong on the form factor. But I also feel like we've kind of solved the wearable computing interface problem — a couple hundred years ago, turns out — and so it's going to take a lot of convincing.
Only real differentiator is maybe the real time translation, but that's not a frequent use case and i think i can take my phone out for that with google translate as needed.
It's too bad, love new hardware, this isn't it for me at least with that price and functionality.
"Hey humane, add a meeting next Tuesday at 2pm'.
"I’m sorry Dave, I'm afraid I can't do that. You have a doctors appointment about your haemeroids"
Why wouldn't I use my existing watch/phone/earbuds/pods instead of paying 600$+subscription for this?
I don't understand the insistence on using voice as the main interaction and ditching the screen.
At least google glass/AR let's me read
I'm skeptical of the usefulness of the hand projection vs a watch. And I think anyone who wants to bring a camera would be far better served by an iphone (or any phone).
Most of the other stuff is just idling. I don't expect I would idle in the same way with an actually good assistant that respects me.
But then I'd prefer an open source Wikipedia/Wikimedia like organisation behind it.
Siri itself is lacking but I expect that to change with an LLM soon.
You can already lock down your phone to prevent distracting apps.
I've skeptical that people would actually choose to go without a phone in favor of this
This requires an always-on device, or always-in-the-cloud server processing your data and pushing updates to your device.
The former is limited by physics (battery), the latter is limited by how much data you want accessed from the cloud. Neither are solved by open source.
It can’t compete in the consumer space, because it doesn’t let you waste time on social media. It can’t compete in the corporate world because it doesn’t have a screen — no email, no spreadsheets, no collaborative chat application we’ve all grown used to. And it can’t even be great for photography, since you need another device to view the photos and videos this thing takes.
If this thing takes off for its impressive AI capabilities, smartphone makers can pump R&D into their AI, and give us this for free as a software update. But right now, the only people who will use this are folks whose job involves scheduling meetings and firing off quick text messages to colleagues and clients.
Incidentally, it was written by Jess Armstrong who later created "Succession".
https://web.archive.org/web/20140208000114/https://kaptureau...
EDIT: had to share promo video https://www.youtube.com/watch?v=arQoSSXKaSQ
People are thinking about the form factor after the cell phone. Apple is busy training everyone to use hand gestures with the new Apple Watch and upcoming Apple Vision. Humane is going down the path of projecting on the hand and touch.
Apple's implementations are for 1 hand operation. You can operate the watch's touch screen while holding a steering wheel for example.
What's the difference between the objectively not great screen that is my hand, and the oled watch that doesn't require both my hands for operation?
EDIT Heck this requires one hand just to see anything. I can look at my watch without any hands!
That presumes there is one. There's not yet a "form factor after the car" for example. Just refinement of the same basic 4-wheeled template, with a few oddball vehicles for niche uses.
A possible indicator here is the apparent lack of demand for small screen phones. To me it suggests that screen real estate is more valuable than portability for most people.
There was a google i/o talk a few years back were they talked about users wanting multi-modal, an example being they ask for restaurant recommendations by voice, then get the list they can view on their device. Both query and results are presented in their easiest modal, and humans will naturally switch between them.
This thing seem dead on arrival. Who wants to hold their hand up like that? Who wants to look at an uneven "screen"? Can you use it while walking or experience the movement in a vehicle? (car, bus, subway)
Is this just a big sunk cost fallacy launch?
It's a combadge.
I repeat: it's a combadge. It solves the self-evident problem of there not being combadges available and in use.
Or, at least, it's almost a combadge. A good qualitative jump forward, but with plenty of unwanted features like subscription (I guess this could work for a Ferengi combadge), screen, wake words, etc. A combadge doesn't need to be an image projector, nor does it need rich tactile controls. But I guess you can improve the product-problem fit by ignoring those features.
Simple example: which way do I go at the next intersection?
Or if I'm driving, GPS is displayed on a giant screen.
Not saying this is for everyone, but there will be users.
For example, the GPS is almost never accurate in Hong Kong, when I visited.
While it looks like there are a few videos of apparent actual demos, I haven't seen one yet where the device (and more importantly, the recording camera's settings) are controlled by an impartial reviewer, and I'm extremely sceptical that this is usable in the real world. There's a demo by the founder where one of the inputs is to tilt your palm up, and even in the demo the projection struggles to compete with the indoor lights, nevermind the sun https://youtu.be/CwSeUV3RaIA?t=205.
The pitch of this seems to be "no more distracting screens, and no need to download and manage lots of apps and services". Except there is a (very poor) screen, it's your hand. And you're limited to just one service and set of apps, the one that comes with the device.
It's all well and good saying that the AI can do everything you want, but the real world (sadly) has copyright restrictions and content licensing agreements which an out-of-the-box service by a legit company will have to abide by. If the song I want to listen to isn't available on whatever music service this product is partnered with, could I transfer music files from my computer to this device? There's a lot of use cases like this where you very quickly start to want an actual screen, and actual methods of input more precise and domain-specific than conversational voice commands.
What a weird example. They say they've partnered with Tidal, which would have 999 out of 1000 songs people look for, maybe more.
Unfortunately, "nobody" has music files any more. Spotify forever.
(Of course readers here are the exception.)
The MIT wearable demo from a few years ago which used a similar concept to project an interface in the real world was incredibly compelling, but mostly because it assumed near flawless real world AI object recognition, along with flawless projection onto said items. They'll need to demonstrate this on this particular device, before this becomes remotely interesting. Yes, it's a "detail", but I think for a lot of this kind of tech, demonstrating just how DEEPLY you can go into the interactions is sort of the whole point if they are thinking of replacing the kinds of devices that we depend on.
Wearing glasses when you can just wear nothing seems like a big progress to me.
The only thing that remains is to make it look better imho
We can always discover new physics, or utilise already understood physics to design something more efficient. Like what if your phone was efficient enough that it could work only on the heat given off from your hands and/or ambient light? Sounds far-fetched but I won't be surprised if this is commonplace within the next 50 years.
Also, you can't view the total eclipse in either locations it stated.
A screen would be useful for showing the details of how it misestimated the almond count, and let you adjust them.
> How much sugar is in this?
> A whole dragonfruit contains 7.31 grams of sugar.
100 grams of dragonfruit contains 9.75 grams of sugar.
https://fdc.nal.usda.gov/fdc-app.html#/food-details/2344729/...
A whole dragonfruit weighs closer to 350-600 grams.
https://www.seedsdelmundo.com/blog/average-dragon-fruit-weig...
I'm guessing they mean a combination, so you need to touch AND do something else. But taken literally the gesture option implies they're also always watching.
> if it's ever physically tampered with, it will require service from Humane to restore operation
So it's entirely non-repairable?
--
I also love the "you can shop in the real world" example where they imply the scenario is him going into a physical bookstore (they say "retail") and yell out that you're looking up if it's cheaper online and buying it there.
- too stealable, by people who will not care that a subscription is needed.
- the act of theft will happen violently and close up, not fun.
- it's an easy smallish act of violence, which means the on-ramp to violence is also easy. Not something most people want to invite into their lives.
- "they" (the Committee) will say phones can also be grabbed. But the equation here is different. With a phone there's no hand on your chest, no tearing of clothing, and for a phone thieves know you will try harder to get it back. With this, after the violent taking, the shock value and the relative disposability of the device will stop most from chasing the thief. This will be known subconsciously if not outright, so the "phones are also easy to grab" comparison does not apply.
- the features are already provided by something most everyone has, a smartphone.
- the level of obnoxiousness of the status signaling is off the charts.
- association with AI is not a positive for many people and is stigmatizing (whether the stigma is correct or not).
- built in camera and recording functionality or even the perceived possibility of recording is also stigmatizing and highly antisocial.
- all the voice UX inhibition concerns others have been mentioning.
- [edit, how did I leave this out, but it's just too obvious]: subscription. We. Don't. Want. More. Subscriptions.
On the positive side, the size is nice, it looks good, and reading stuff off your hand is a cool idea, although it will look pretty goofy. But no.
I even doubt there will be much theft of these. People will simply forget these, and stop using them.
So they showed one clipped to a jacket. Don't they take the jacket off? What's the intended usecase? That you take it off and re-attach to various clothign as you dress/undress? It also looks quite heavy, so most T-shirts and other light items of clothing are not really suitable for this.
And they really doubled down on that decision by also embedding it (“pin”) in the name, if not the identity, of the product.
They could have coined a word (I’m not claiming this is not cringe) “pindant” as in a dual use pin-or-pendant item, and bought more flexibility, for example. Edit: somebody already coined that word, see the dot com (sfw), lol.
Clearly very confident in their product.
A product like this makes it very difficult to verify what it is telling you.
As others have pointed out, their own product launch video has several inaccuracies in it.
#include <iostream>
int main() {
std::cout << "Hello world";
return 0;
}
Seems like you can determine for some programs what they will do.I'm especially excited about the fact that they found a really low-barrier user interface for using CV and AR-type functionality -- like, without having to put on silly glasses, and without having to use a second device with a screen in addition to the pin.
Come on, this is cool! Or would you have designed (and built!) a better device?
But why only focus on some potential shortcomings instead of appreciating the positive aspects?
Btw, there is no tech device out there for which I couldn't come up with a list of critical questions like yours.
I'm limiting my points to the physical usage concerns, there are more concerns if we broaden the context, many other commenters have pointed them out. This is not even the full list of physical concerns. What people wear will have a big impact too
I also find it curious that a former Apple exec formed this company. I'd assume Apple itself would want to pursue this internally, as such a device would be yet another killer addition to the iron grip of the Apple ecosystem.
It's nice to see this product isn't actually vapor. Congrats to them.
I have a similar feeling about augmented reality glasses.
It was/is an amazing experience. It's really a hardware miniaturization at this point, except that M$ canned the device and team to focus on other things. Really thought this was their opportunity to build a device that would dominate the market
It seems like a shame, Apple is entering that market soon with a device that sounds like it’ll be the same price, and… I dunno, I’d expect the real vision advantage to be a pretty strong selling point.
Fortunately, much like the HashiCorp hoopla, there is a group taking part of the project forward in an independent org. I'm looking to deploy an MRTK demo to the Q3, hoping Immersed will pick it up for the next iteration of their app
The watch is an accessory to the phone that adds features plus offers convenience and if you want to, it can temporarily substitute your phone like when you go to the gym. A cellular Apple Watch combined with Air Pods can do a lot. And both the Apple Watch and the Air Pods have use cases in addition to that in other situations. I don't see that here at all. I see a device with a very limited feature set.
Edit: Wording
https://www.theverge.com/2023/11/9/23953901/humane-ai-pin-la...
This thing is incredible and will eventually crush the iPhone. Solves iPhone addiction while retaining the utility of an iPhone? Solid gold.
In reality, they want to read news while waiting at a doctor's office, play games while they take the subway, and see Instagram updates from friends throughout the day.
And if you already want a less capable device, it's called an Apple Watch, but it comes with a little screen that is way more useful than laser projection, and will soon surely have a powerful LLM it can access. (And paired with AirPods it does a much better job preserving your audio privacy.)
So it's hard to see how this is going to succeed, when Apple can just copy the good part (LLM) as part of the Watch.
Most people saw the utility and the use cases of Dropbox even when it launched.
What's the utility and use case of this? What problem does it solve?
It's just a smartphone, except you can't run third-party software, can't directly interface with it, and can't connect it to other machines. And instead of holding an N-million pixel, M-million-colour, extremely high-constrast display directly in your hand, you have to indirectly project (meaning extremely LOW contrast) a single-colour display onto your hand from a projector that's shaking around being clipped to your clothes.
The only single hypothetical upside I can see to this tech is that it might lower the two-second delay in looking at my phone caused by putting my hand in my pocket before raising my hand, but you could say that that goes against the goal of solving phone addiction.
This is not a thing. "Screens" aren't 'separating us from one another', or 'distracting us'; that's fuzzy verbalistic nonsense, made up by marketers who want to sell you non-phones, and bloviating op-ed columnists who don't have a clue. It's so ridiculous, that everyone has seen the memes debunking it.[0][1]
True invasiveness is expressed as: "how long does it take me to do this thing I want to do?" In other words, you need a human-computer interface that reduces friction as close to zero as possible. The phone won because it's the best at that. The "pin" is orders of magnitude worse, so it won't catch on.
[0]: https://xkcd.com/610/ [1]: https://imgflip.com/i/1swr7j
You know what, I just don't think I'm the target audience for whatever this site is trying to sell.
Talking about alternatives often leads to mere concern and agreement without action. Presenting an actual alternative, however, deserves respect.
Yet, there's a hint of skepticism in my appreciation. Why the 'AI' pin? The constant mention of 'AI' arouses suspicion about the product, recalling a time when 'AI' was not a part of their lexicon.
Nonetheless, I wish them good luck with "AI" pin.
This thing has no business being $700.
I guess they really do hate smartphones…
They are making a device to replace the screen, yet the screen is a rich source of information! People love screens! Screens can present information non-linearly and also interactively.
While voice and sound is always linear.
And yes, this thing has a display, but it's low fidelity and cumbersome.
On another note, this reminds me a lot of the short story The Perfect Match by Ken Liu. The story isn't ground breaking but is worth a read and harps on AI assistants making decisions for people and driving biases based on the corporate agenda and sponsors (not to get too tinfoil hatty).
The pin form-factor is awkward. At least with a watch, you have watch functionality to fall back on, making it immediately useful, and you can discover incremental functionality--health, message, alerts, etc. This is all or nothing (and I think it's going to land closer to nothing).
Get it to the point where it can constantly observe my surroundings and make the sort of suggestions a partner might ("if we stop at the hardware store first, we can get those fresh bagels Bob likes from the place that closes early", ) and maybe there's something to talk about.
You just might have to deal with reactions to always on cameras and the annoyance of being admonished by LapelClippy on a regular basis
Siri question type things I can use my watch for real time translation is neat but it has to be a big and frequent use case to differentiate from just taking my phone out and using google translate (which i believe i can also do on my apple watch).
Also agree this is more a smartwatch competitor than a phone competitor. The fact that smartwatches sell at all is proof there are (much smaller) markets for wearable devices that do stuff that can be done at least as well on a phone if you get it out your pocket, the argument for separating the powerful internet connected functionality from the watch and having it on some other wearable on your chest is that actually I like the wrist-mounted device that tracks my activity and sleep to not need charging every day...
I mean, has any skater ever thought "I want to listen to a hip hop playlist that mentions skating" - as someone into extreme sports I highly doubt this has ever been on a skaters mind.
And how many people are walking around thinking they want a picture of a solar eclipse?!
The device and its marketed use cases just seem so ridiculous, it honestly blows my mind that this thing exists.
I think the Meta Ray-Ban glasses are probably a better concept even without a screen just because of the better placement of the audio and camera. But I'm glad people are trying different things.
LOL. Not made for this planet. Heck, put a jacket over it and your body heat could take it over 35.
I salute efforts to develop alternative computing devices, interfaces, and ways to interact.
I have been using my Apple Watch in an ‘only device carried with me’ mode, unless I will be wanting to take pictures or read an eBook when I am out of my house. As someone who has been using computers since about 1964, it is so refreshing to just have minimal connectivity - this helps being more present in the world.
I would love to be able to fast-forward and see how devices like the AI Pin do commercially in the next 5 years. How many people are like me and want to digitally disconnect, except for communicating with family or close friends? I would bet we are in a small minority.
Not having any third party apps is painful. No audio books, or podcasts? I get those on my Apple Watch used with AirPods. Anyway, I wish this company well!
It’s almost as if they held their two most apathetic employees at gunpoint in a laboratory. Safe to say I’ve never seen that level of ennui in any startup’s presentation before.
Also, I like how the focus appears to be on how it will benefit the user, instead of focusing on tech specs.
Data privacy and personal security aside, I understand there will be a reciprocal action between lifestyle and technology. We might have to change our lifestyle to make room for and benefit from new technology tools, just as we have for smart phones.
If kids get interested in it, then it will have a chance. If it's cool and fun, then it has a chance. If it actually makes things easier and better, then it could take off.
But if it has that cringe factor like google glass had, then it will never get anywhere.
The possibilities are awesome, but something like this requires a reinforcing feedback loop on top of a network effect to become successful.
You will not separate people from video calling and media consumption, sorry.
But imagine you had this as a peripheral for your phone. You could be on a call with someone, face to face, and then say, "hey, I will take you for a walk". Switch to the pin's camera and keep talking while your friend now sees more or less what you see as you walk around.
Once the AI Pin replies with a "sorry, couldn't get that" a few times, people will give up on it and reach for their phones. I could see it finding some success in the accessibility market, but outside of niche applications, I don't think this thing sticks around.
As I sit in a crowded coffee shop at a shared table, with three people who are on meetings and talking away about sensitive things. Right next to one another! Along the wall there are two people chatting on their phones to family, one of them on speaker.
People don't care, they don't have manners.
A coffee shop is NOT supposed to be quiet. If you want quietness go to a library.
So far I can do that with RAG using my tweets and articles, but it's not sufficient; I need to incorporate my ongoing real-life experiences into the model. This relies on advancements in battery technology and compression algorithms.
Edit: wow their tech specs are actually detailed. 125 degree FOV.
The best use case i can see for a device like this a hands free recording and support for law enforcement, rescue and emergency teams.
Imagine a paramedic recording all the info about the victim of a car accident and being able to project relevant information or questions on the asphalt or any other surface at night. The rescue crew would have both hands free to stabilize the victim, maybe even get the victim's medical history without having to stop everything to type on a smartphone or tablet.
I know a handful of folks that would love to have a replacement for their phone - something less bulky that they can carry around in place of their phone, that still accomplishes most things a phone would. Most recently I heard of people trying this with the Apple Watch Ultra, and ultimately giving up on it. This Humane pin seems even less capable?
It seems like a great attempt though. I'm excited to see some new innovation in this space!
There has been a shortage of ADHD medication because people are taking it as an enhancement drug (tech folks do it a lot) and people who need it aren't getting it. My wife can't get hers, it sucks.
I told her 'I guess everyone in tech is also short on ADHD medication now'. ;-)
But I agree with your point that some combination of Apple Watch, AirPods, Meta AR could provide a better experience and probably future AR glasses will have better display technologies.
I wish though pico projectors got more maintstream in devices such as laptops, tablets - there could be many useful applications for indie devs.
more muggable than a "I heart NYC" t-shirt
But then Google already has access to all my data - via Gmail and my phone (android)
Humane clearly appears to be a category creator in the making
However those are things that can change. So I don't worry about that.
The form factor is innovative, and the laser display (which they didn't lean into) is very cool!
“When is the next solar eclipse and where is the best place to watch it?” Returns actually relevant and accurate results from the web.
“How many grams of protein are in a handful of almonds?” Returns “7” with a citation link which is closer to being correct. Hard to compare apples to apples on this one since Siri doesn’t have a camera, but the Pin demo is very obviously wrong.
Starting at 699$ damn that's pricey considering phones are already stolen out of people's hands/off tables in busy cities, do they think a thief won't rip this off someone's shirt if they see it?
The display: pretty cool, but tracking seems a bit wobbly, I haven't seen a demo of it in bright sunlight and idk if it would do that well.
Chest mounted camera: I mean being able to hold up something and ask it questions is pretty cool, like identifying something, or getting the price of an item from a store vs online to make sure you're paying a good price, etc. Would be great for travelling.
The translation: maybe just demo but it's too slow and it doesn't show what happens when you walk through a crowded market or airport with tonnes of people speaking in diff languages, you bump into someone behind you and they say something in Spanish. How does it know to translate? You would, because you felt the bump, but the device didn't.
Hand gestures: these always suck and never work well. You always end up doing things you didn't intend unless you interact with it unnaturally/according to its quirks.
Asking it questions etc: our phones do this already with Siri/Google ass/Bixby. And I'm sure plenty upcoming as LLMs/elevenlabs style voices are used.
"Make me sound more excited" dystopian lmao.
"Catch me up": "Jane said to remember to get the report done for next week. Thomas and Aaron won't to know if you're still up for the movie tonight. BigBuns27 said thanks for railing me so hard last night big boi" yeah I want that potentially coming out at any random time during the day.
My conclusion: this is fluff, we already have powerhouses with long ass battery life in our pockets already. The real side of this is an accessory that lets us interact with our phones faster and more naturally. A chest mounted camera is fine but still not as good as glasses/contacts which can see what we see, very important because we'll usually be looking at something. Glasses/contacts can also provide silent HUD help, they're much easier to provide consistent display experience vs projecting from a moving torso to a moving hand.
I may be biased as someone who already wears glasses but damn was I sad when Google Glass got ridiculed (mostly by older gens) as if they'd continued & an improved version was out today I would 100% own one and develop whatever I could for it.
Being able to look at something in another language and have translations superimposed over, be walking around an unfamiliar place and have subtle direction hints provided, would be amazing. Combine with bone conduction speaker and some sort of subvocalisation system to ask questions/give commands and it would be absolute fire.
I don't want this weird kickstarter-esque device; I want Google Glass back ;~;
> A Buddhist monk named Brother Spirit led them to Humane. Mr. Chaudhri and Ms. Bongiorno had developed concepts for two A.I. products: a women’s health device and the pin. Brother Spirit, whom they met through their acupuncturist, recommended that they shared the ideas with his friend, Marc Benioff, the founder of Salesforce.
> Sitting beneath a palm tree on a cliff above the ocean at Mr. Benioff’s Hawaiian home in 2018, they explained both devices. “This one,” Mr. Benioff said, pointing at the Ai Pin, as dolphins breached the surf below, “is huge.”
> “It’s going to be a massive company,” he added.
This product was also named a "best invention of 2023" by TIME magazine before it was even released. Entirely by coincidence, Marc Benioff happens to own TIME magazine.
HBO's Silicon Valley may be over, but the real world Silicon Valley is still going stronger than ever.
Almost the entire beginning of the video is about which colors are available and how the battery snaps, with zero hints about why I would need a cringe projector on me.
I can't believe this was shipped by ex-Apple people. Imagine Steve Jobs introducing the iPhone like this: "We are introducing a revolutionary new device. The first thing you should know about it is that it has a charger and an Apple processor. The second most important thing: here is how the battery works."
I couldn't figure out why he kept touching and adjusting the pin on her chest, a thing I would never do with a coworker. All I knew was that she was CEO and he was Chairman, so I knew it was a joint decision. This makes so much more sense.
Sign no one internally is being honest with them, feels they can say "It's bad"...
> Altman owns “14.93% equity and voting through a number of holding companies none of which individually holds 10% or greater ownership interest in Humane,” the filing states.
https://www.lowpass.cc/p/humane-ai-pin-cellular-mvno-sam-alt...
He was fairly outspoken about getting (even more) rich off of other investments and believing OpenAI was simply too important to make it a conflict of interest, and mostly considers it a nuisance/distraction. That's fairly arrogant, and, again, might be completely off but I still do believe he means that and I would give it good odds to be the entirely right course of action, if most impact/most quickly is what you are going for.