AirPods as a Platform
julian.digital
julian.digital
Most of the time, if I'm wearing headphones, it's so as to not disturb others around me. Otherwise I'd play it out loud.
This benefit goes away when everyone around me suddenly hears me bark out loud to adjust the volume, or send a text, or what-have-you.
I'm a big believer in audio-as-a-platform (particularly the AR possibilities), but I hate audibly trying to speak to a computer. It's by far the worst input interface.
(On the other hand, much like cameras, the best interface is the one you have with you. I yell at my Google speakers all the damn time, because my hands are busy around the house. But those are also speakers, not headphones, and therefor a different use case.)
Maybe that's the key to its network effect as a platform. If you don't want to hear everyone talking to themselves in the library, you'll need noise canceling headphones too.
Other than that though, I haven't found any practical use for them. Maybe once AI improves a lot, and it feels like you're talking to an actual person, it might be more useful.
(And then said list is shared with my family.)
In my eyes the primary thing that's standing in the way of voice assistants being useful isn't even as high of a bar as general AI, but just the ability to parse a command into multiple, potentially chained commands. Even with the inability to figure out things like context that would boost usability a lot — for example, it'd allow commands like, "set timers for 5 minutes, 10 minutes, and an hour" instead of having to make each request separately.
/S
The first time I tried the office kitchen microwave, I had to ask someone from accounting standing around, how to heat my cup, because that stupid thing just responded with a condescending beep to whatever I pressed :/
I also have a five year old who is pretty competent at talking to Alexa even though she can’t intuit what the button controls mean on the microwave oven.
What a lot of these natural language systems do is they force me to intuit their interfaces instead of just looking them up. It's more complex, harder to learn, and much less powerful. What I want is very reliable, accurate voice detection with a well-designed, composable interface.
I want to treat Alexa like a computer, not like a person. People are not the most convenient interfaces to interact with, computers are better to interact with for me than people are.
I understand that different people are in different positions, some people want to have a conversation, but you're not going to make a voice assistant that's good for me if you follow that goal. At some point the ecosystem needs to fracture and diverge so that normal people can use whatever NLP interface Google/Amazon/Apple spits out of their AI opaque-boxes, and people like me can use a voice interface that is designed around well-tested decades-old computer UX principles that have been proven to work well for power users.
My vision of a voice-operated utopia isn't treating Siri like a person, it's on-the-fly composing a complicated new task by voice that is saved for later use. It's using timers as a trigger for other tasks with some kind of pipe command so I can tell Siri to send an email after ten minutes, or so I can have Siri look up a search and seamlessly pipe the result into a some kind of digital notebook.
Saying "hey Siri" is fine if I'm in bed or in the shower, I don't need quick access to a shell in those places necessarily. That's fine to have as a backup. But for normal operation, if I'm wearing a smartwatch, it will pretty much always be more convenient and faster for me to tap and hold on that watchface than it will be for me to say "hey Siri".
I mean, that's a boring answer, but there's also a reason why my computers have buttons. I wouldn't want to use my phone because that's in another room or in my pocket. But a watch will always be reachable in less than a second, and the modern watches are waterproof, and I don't need to look at anything to use it -- I can just tap my watchface and start talking. And if my hands are dirty, or I'm carrying groceries, or I'm in bed, falling back to "hey Siri" isn't the end of the world in those scenarios.
In practice, when I see people interact with voice assistants today, they stop what they're doing, they give the command, they listen for a confirmation, and then they start what they're doing again. The biggest bottleneck there for their speed is precision -- they intuitively know that they need to stop what they're doing and optimize for the device. The precision, and the delays that are built into the UX to confirm what's happening -- that's the bottleneck. So if there's an operating mode that is just as fast and way more precise, we should just do that, we don't need to use voice triggers 100% of the time.
Bonus points if we're wasting processing time for a voice assistant to make a round trip and process the audio clip to try and figure out who's speaking. The person who pressed their watch is speaking, boom, we can get rid of that response delay now. How much time are we wasting trying to come up with wake words that optimize for both speed and precision -- when using wake words only as a fallback would allow us to make them more precise because they could be longer, more deliberate phrases?
Same here. Just wish it could handle setting multiple timers (yes I realize you can set multiple alarms, but those aren't one-time use).
Hey Siri, turn out all of the lights.
Hey Siri, play Ocean from Ambient Sounds.
Not a huge deal, but enough of a bump to break the flow.
I've got one for "Siri, good night" that turns off the lights, sets the phone to DND, turns down the brightness and starts Sleep Cycle in the correct mode depending on the day (alarm on/off).
They also support custom "routines" that you can program. We have a Clever Dripper [1] and use Tom's (of Sweet Maria's) recipe: stir after 90 seconds, drip after 4 minutes (2:30 after the stir). So I created a routine called "coffee timer" that triggers those two timers, and when we pour the hot water into the Clever Dripper we can just say "Hey Google, coffee timer" and it sets both of them.
Actually I put in four variations so we don't have to get the language perfect:
"coffee timer"
"coffee timers"
"set coffee timer"
"set coffee timers"
[1] https://www.sweetmarias.com/clever-coffee-dripper-large.html
Unfortunately iOS only supports one timer at a time so I can't make tea while my car is plugged in.
"timer 1 hour 30" is parsed into "Make an alarm at 01:30 named Timer."
Every time I do laundry I get woken up at stupid o'clock the next morning haha.
Edit: Sorry to those below for the confusion I seeded. To clarify I mean the 1 hour 30 is the bit that doesn't work. If I add minutes to the end of it it works perfectly fine.
That's how I talk. Thats how I've talked for 32-$childhood years.
I'm not having a conversation with the thing, I want it do do something. Command, parameters.
computer, list the invisible files in my home folder
but in a shocking turn of events, when you use it wrong it doesn't workIn what world does "timer one hour 30" parse into "Set an alarm for 1:30am and call it timer"
Presumably if a French person had trouble you'd tell them to just speak English?
Parsing time related stuff in general seems to be an issue. "Set a timer for a minute fifteen" makes an alarm named "Timer" set for 3 PM.
But if you say "Set a timer for a minute fifteen seconds" it works fine.
Curious if anyone else can duplicate the "a minute fifteen == 3 o'clock" issue or if it's somehow hearing me wrong.
Other simple time-based requests such as "what's the time difference to singapore?" don't work on siri either, which is irritating as I work across multiple time zones and I'm forever figuring out time differences.
* Hey siri what’s the weather
* Hey siri add broccoli to my grocery list
* Hey siri remind me when I get home to take out the recycling
Same principle. Simple request, hard to get wrong, low stakes, faster than clicking
* Hey siri remind me every other Tuesday at 8pm when I'm at home to take out the recycling
Some things are more easy to say than to configure with 20+ taps on the Reminders app.
I have huge gripes with "what's the weather". Siri will for sure say the temperature, which means nothing in a city where the windchill is often ten degrees lower.
Are you using a specific app for this or is it just a list in Notes called 'grocery list'?
Works beautifully for me and the lists are shared with our household so it updates for everyone.
1. Hey Siri, add Olive Oil to the grocery list. Super easy. When I’m cooking and running low on something, no need to break my stride and pull out my phone, or clean my hands. It’s immediately into my grocery list and out of my mental inbox.
2. Hey Siri, play some Jazz. Or whatever music. This is nice and easy to get some music on the HomePod for either cooking, working, or dinner background. The only annoyance is that Siri can be super particular at times unlike when searching Apple Music on my phone. Also sometimes my kids hijack my music selection with their own, hehe.
I also have a few automations that are set up for when I leave the house to make sure music turns off and my thermostat is set to "Away" mode.
I've generally really enjoyed basic smart home stuff. I have a handful of smart wall switches that work well, and for the lamp a wall-wart that I use. One of these days I'll just retrofit some actual lighting in the living room and can skip the lamp situation but until then... this is a nice workaround.
"OK, playing Supertramp."
Somewhere along the way that became my normal Siri experience, and it wasn't always like that.
* Call my wife (or $person)
* tell $person $message - sends a text eg: “tell Dave ETA 25min”
* ping $myname iPad - plays sound on my iPad so I can find it.
* set timer for $duration (I wish Siri wouldn’t be so loud in acknowledging)
* remind me $task on $date $time (you can switch the arguments around)
* take me home
* redial
* take me to $location (mostly used in my car)
Edit: found a list online [1] - seems to cover a lot of default commands
Hey Siri, speak quieter / speak at 25%
Example (NSFW/NSFL/gross): https://www.reddit.com/r/MakeMeSuffer/comments/l8d0qo/who_kn...
1. Timers. Super useful when cooking. Also my kids can use it as they are doing remote learning and need to know when to get back to their video calls.
2. Playing videos while cooking. Sometimes I enjoy watching a sitcom while cooking on my Echo Show. Or I might want a cooking video though that’s more rare. My roommate uses it all the time for questions like “how long do you bake salmon?” With mixed results.
3. Controlling lights. “Alexa give me a light” as I walk into a room is way easier than turning on three separate switches in separate parts of the room. Turning them off is equally nice.
4. “Alexa tell me a kids joke” is a frequent thing we use.
5. Answers to random questions while at the dinner table. “Alexa who is the prime minister of New Zealand?” That type of stuff. It feels more natural than whipping out our phones.
6. We tried using Echo Show’s drop in feature but it is just too intrusive as compared to something like FaceTime. The other side doesn’t have to pick up the call. You just are in their house, camera on and all.
7. My kids really like when we play Harry Potter quiz. It’s a silly app but it is somewhat entertaining.
8. Really funny “routines” (their word for scripts). “Alexa, set condition two throughout the ship” to turn off all the lights. “Alexa, release the kraken” to set off the Roomba, etc.
9. My kids listen to podcasts all the time.
10. My kids use it to help them spell difficult words.
What I wish I could do a bit more with it is integrate it with things like random status dashboards. I combined a power metering AC plug with my washing machine and Home Assistant to know when it’s done running. Would be nice to be able to say “Alexa notify [roommate’s name] when the washer is done.”
Overall I think something much simpler that does processing locally could replace it for me but so far these things are cheap enough (echo dot) to put in every room.
It's a lot easier to say "Siri, play podcasts" (Which triggers a Shortcut to start playing a specific playlist in Overcast) rather than opening my jacket, digging out the phone, taking off my gloves and fumbling with it to get the podcasts running.
That would definitely be something I would buy on to.
I'm imagining a system where you can just use nods/head shakes to move through some sort of binary decision tree to execute some basic interactions, like reacting to incoming alerts/messages.
new text, music pauses, siri reads it to you: "Your mother asks if you'll be home by dinner, Would you like to respond?" shake head no -> interaction cancelled, music resumes nod head yes -> "Ok, how would you like me to respond? Yes, you will, or no, you won't?" gesture head for appropriate response
Easy? Dumb/ridiculous? Sure, you can't get suuuuper deep with the decision tree and it's tough for non-binary responses, but it's enough of an interface to have a meaningful, non-verbal engagement with a computer.
Just like with the M1, if you squint, Apple is testing and iterating in the open. The spatial audio is a good example but so is the Watch’s auto-detected hand washing countdown.
There are public sprinkles of this coming platform elsewhere, such as in Apple fitness workout HUD Rings widget. The proximity-based handoff is another.
The author is correct that Siri is not a good platform but for reasons they do not identify.
Voice based interaction model is weak from a UX perspective. But for Apple it is even weaker because the company is unable to use any of the unique advantages it holds over the competitors.
For example, apple’s array of services, control over the technology stack, reliable and secure intra-device communication, the App Store, the iPhone as a unified configurator, access point, update manager and biometric authenticator.
You can’t pull that stuff out of a hat.
Siri sucks. I have a few HomePods, use plenty of homekit and try to get the most out of it. But it is bad at almost everything it sets out to do.
Siri is clearly not the focus for the company, and if anything it sent competitors scrambling to own a space Apple doesn’t even want.
The interaction model includes physical hardware, like the big crown on the new APMs, but I suspect it is likely going to be based largely on eye movement. Something not too twitchy.
The enormous amount of sensor data from watch and AirPods are like the gps and gyroscope of iPhone. Apps can require either or none.
So I think the author is right that AirPods are important but they are not the center. They are a component of the next platform.
When golfing, I want to keep track of how far I hit the ball, the club I used, and where I landed. There are apps for this, of course, but I can't use them. By the time I've arrived at my ball, I'm not going to stop, take out my phone, and start fidling with UI controls to select a club or confirm a location.
I would love to have an app that let me keep one AirPod in my ear, and allow me to track my golf game. The UX would be something like this:
1. Arrive at course, and use phone to select the tees and confirm the course I'm playing. Start the round.
2. Tap my AirPod and say, "Teeing off on hole 1 using driver"
3. Hit the ball
4. When I arrive at my ball, tap again and say, "hitting seven iron"
5. When I sink a putt, tap and say, "next hole".
From just those interactions, the app could keep track of every shot, and also keep my score and number of putts. I could choose to not announce each and every shot if I wanted to, and instead say, "add three strokes" once I'm done with the hole.
I could also ask, "How far to the middle of the green?" and get a distance in my ear. "What did I hit last time on this hole, for this shot?" (Answer: "You used a nine iron, and hit it 107 yards")
All that would be killer for me. Nicer than staring at my phone screen in the sunlight, and looking like I'm farting around to the players waiting for me to clear the fairway.
Anything like this exist today?
[Blog post announcing the feature](https://udisc.com/blog/post/picture-this-udisc-unveils-map-s...)
An alternative solution would be a smartwatch app which would spare you chanting out on a course.
You might also want to check these: https://www.wareable.com/golf/best-golf-wearables-gps-watche...
Their only real value proposition is basic interactions when the user's hands and/or eyes are busy with something else.
Until they can get smart or fluent enough that they can rival the effectiveness and accuracy of hands-on-screen interaction, the use cases will remain fairly niche.
More programmable, multi-modal interaction is a step in the right direction, but it'll require a lot more.
"Hey Siri, what's the weather today?" ==> "Hey Siri, weather"
"Hey Siri, set a timer for 10 minutes" ==> "Hey Siri, 10 minute timer"
Transparency mode is a critical success even if it’s so boring as to be unremarked upon by most people. An AR device is most useful if it’s ubiquitously available. Transparency mode makes that possible (even if it could use improvement). The device also needs to avoid a negative social stigma. Airpods have largely achieved that as well (at least along younger people). That is partially marketing/brand image - but it is also based on utility. Older folks would generally find it rude to leave headphones in while having a conversation because they assume the listener isn’t listening, but I work with a lot of teens, and it seems like they couldn’t care less. It’s understood that the speaker can still be heard. A quick tap/squeeze is the social signal.
The author is right that there is huge untapped potential in auditory augmentations, but the focus on verbal input control is misplaced. It’s simply too obtrusive for public environments.
If I were betting, I’d say Apple won’t open up this kind of functionally until the (cross device) input control is generally codified, and that scheme will be intrinsically linked to a forward facing camera/sensor package to provide contextual awareness and implied user attention & intentions (i.e. glasses or similar).
Working on AR UX would be incredibly exciting.
It does leave room for improvement. Perhaps some visual signal that shows that you can hear your surroundings clearly would be better.
- Not entirely true. It's the pairing of the device with an iOS or capable watchOS device.
"Why has no one thought about additional buttons or click mechanisms that allow users to interact with the actual content?"
- It's called a smartwatch (The pebble was really nice at this). or generically bluetooth radio controls.
I wish these design analyses talked about the material input costs needed to produce the thing we might perceive as a 'platform'. I just see more batteries, wear, expense, etc.
For me personally there is a suite of tools involving audio books and note taking that would change my life: A remote with a few physical buttons to rewind, switch to record-mode, skip sections. Speech to text with full text search. Voice recordings tied to what I’m listening to. Basically, I want to be able to work through a difficult audiobook while walking around.
Microsoft added PowerPoint forward/back control to their earbuds.
https://www.businessinsider.com/microsoft-surface-earbuds-pr...
You'd no longer be sure if someone is having a seizure or trying to stop the podcast she is listening to.
That's an interesting perspective. Theoretically, the best way to get rid of hardware is to move as much as possible to "the cloud"[1], yet Apple isn't very good at cloud. (At least, not as good as Google and Amazon.)
So let's say we're headed to a future where the only physical electronics anyone owns are wearables: watch, glasses, ear buds. No phones, tablets, laptops, or desktops. Just wearables.
In that scenario, who wins? Apple is best suited for making that hardware (by a long mile), but Google and/or Amazon are better suited for handling the software in the cloud.
I'd place my bets on Apple catching up on cloud faster than Google or Amazon catching up on hardware.
However, if we took it a step further and went full Mana[2], tapping right into the nervous system, my bet would be on Google winning that one. They have the cloud capabilities and expertise, but Alphabet also has some experience in health and biology (if I'm not mistaken).
--
[1] I know, I know. "Cloud" is just someone else's computer. It's also more than that.
That's an interesting perspective
It's Steve Jobs' perspective. He talked repeatedly about technology disappearing into the background, and one day we would have technology so good that we wouldn't even see it. It would disappear into the walls.
To me, it's the ultimate expression of making computers work for us, not the other way around, which is mostly what we have now.
Rather, I'm saying they're nowhere near as good at it as Amazon or Google, and I anticipate that this gap is only going to grow.
An alternative view is that Apple’s biggest product puts enough compute in your pocket to make “cloud” unnecessary in a lot of cases. My phone is somewhere between a t3.medium and a t3.2xlarge (based on ram and cpu cores respectively). That can provide a lot of local compute for my wearables. And those wearables are gonna need network anyway, so either that all end up with 5G cellular radios (and 4/3G fallback) and enough battery to run that, or one device provides the network hub and the tiny things in your ears and the glasses sitting on your nose can have lower power radios and smaller batteries.
I reckon watches/glasses/earbuds(/cars/tvs/etc) all relying on a phone is a reasonably sensible model, rather than each of those devices having completely stand alone capabilities.
(And, the idea of Google tapping into my nervous system??? No thanks... I’ll proudly be a data center smashing neo-Luddite before that happens to me...)
Apple isn't very fond of cloud, which a lot of people appreciate.
For example, if I am playing music on my Airpods from my iPhone and my phone is in my pocket, turning the crown on my apple watch is a really neat and intuitive way to change the volume. The first time I did it and it worked it felt like magic.
Similarly walking down a street and getting audio directions on AirPods almost works - but if that's combined with a small map on my watch it works much better - better than a phone.
But at the same time, neither a watch or Airpods are going to be the right way to send a private text message on a quiet bus, and because a giant new unifying technology isn't with us yet, I suspect a hybrid approach is going to be with us for a while.
Two commands (and that's if you've even got both in) is not enough!
Copyrighting this right here btw.
I'll take 10% of all future sales please.
Edit: Removed the points that he already addressed after seeing comments below.
> The input mechanism I describe doesn’t have to be a physical button. In fact, gesture-based inputs might be even more convenient. If AirPods had built-in accelerometers, users could interact with audio content by nodding or shaking their heads. Radar-based sensors like Google’s Motion Sense could also create an interesting new interaction language for audio content.
> You could also think about the Apple Watch as the main input device. In contrast to the AirPods, Apple opened the Watch for developers from the start, but it hasn’t really seen much success as a platform. Perhaps a combination of Watch and AirPods has a better chance of creating an ecosystem with its own unique applications?