Apple previews Live Speech, Personal Voice, and more new accessibility features
apple.com
apple.com
I think there is an important little nod to the future in this announcement. "Personal Voice" is training (likely fine tuning) and then running a local AI model to generate the user's voice. This is a sneak peak of the future of Apple.
They are in the unique position to enable local AI tools, such as assistance or text generation, without the difficulties and privacy concerns with the cloud. Apple silicone is primed for local AI models with its GPU, Neural Cores and the unified memory architecture.
I suspect that Apple is about to surprise everyone with what they do next. I'm very excited to see where they go with the M3, and what they release for both users and developers looking to harness the progress made in AI but locally.
"Hey Siri, I had a meeting last summer in New York about project X, could you bring up all relevant documents and give me a brief summary of what we discussed and decisions we made. Oh and while you're at it, we ate at an awesome restaurant that evening, can you book a table for me for our meeting next week."
Just last night, we were entertaining our toddler with animal sounds. It worked with “Hey Siri, what does a goat sound like?”, then we were able to do horse, cow, sheep, boar, and it somehow got tripped up on pig, for which it responded with the Wikipedia entry and told us to look at the phone for more info.
"Hey Siri, whats the weather?" and "Hey Siri, what the X-day forecast". Everything else is a huge mess.
But you know what? Still better than Alexa for managing my smart home stuff. By miles and miles, IMO.
I have thousands of contacts, lots of photos, videos, and emails, all in Apple’s first-party apps and yet Siri is more likely to respond with a popular song or listing of news articles that’s only tangentially connected to my request.
This becomes more complicated when Siri is the interface on a homepod in a shared area. Who's data and preferences should be used? Ideally it would recognise different voices and give that person's data priority, but how much can/should be shared between users? Where are these data - they shouldn't be in the homepod, so it would have to task the phone with finding the answer. I'm sure something good could be done here, but it wouldn't be easy.
Doesn’t the locally-run LLM bit solve this particular grievance? I’m picturing a personal AI à la Kim Stanley Robinson’s Mars trilogy.
Well, this is about adding ChatGPT-level smartness to Siri, not just the semi-dumb assistant of yore.
https://www.macstories.net/ios/introducing-s-gpt-a-shortcut-...
Some examples from that blog post:
> I’m feeling nostalgic. Make me a playlist with 25 mellow indie rock songs released between 2000 and 2010 and sort them by release year, from oldest to most recent.
This doesn't just return a list of songs, it will create the playlist for you in Music.
> Check the paragraphs of text in my clipboard for grammar mistakes. Provide a list of mistakes, annotate them, and offer suggestions for fixes.
> Summarize the text in my clipboard
> Go back to the original text and translate it into Italian
I haven't tried it myself, but it has other integrations like "live text" where your phone can pull text out of an image and then could send that to GPT to be summarized.
Version 1.0.2 makes improvements for using it via Siri including on HomePod.
siri: The temperature is currently 54 degrees. Expect clear skies starting in the evening, going down to 52 degrees tonight.
me: hey siri what's the high today
siri: the high today will be 84 degrees
me: hey siri will it rain today?
siri: Expect heavy thunderstorms around noon
Note nothing it said in the original was actually wrong...
Meanwhile, Google and Amazon have decided that the data center costs of their approach just aren't worth it.
>Google Assistant has never made money. The hardware is sold at cost, it doesn't have ads, and nobody pays a monthly fee to use the Assistant. There's also the significant server cost to process all those voice commands, though some newer devices have moved to on-device processing in a stealthy cost-cutting move. The Assistant's biggest competitor, Amazon Alexa, is in the same boat and loses $10 billion a year.
https://arstechnica.com/gadgets/2023/03/google-assistant-mig...
Both companies made big cuts to the teams running their voice assistant tech.
As someone who dictates more than half of their messages and is an incredibly heavy user of Siri for performing basic tasks I really noticed this sudden decline in quality and it's never got back up there - in fact, iOS 16 really struggles with many basic words. Before iOS 13. I would have been able to dictate these two paragraphs likely without any errors however, I've just had to edit them in five places.
https://machinelearning.apple.com/research/recognizing-peopl...
Let me sync my information to something local.
I know there's Nextcloud, but it's not as seamless as iCloud.
saying that, their intentions were good, I'm always horrifically amazed at the number of cookies used whenever I see the preferences popup. I honestly had no idea how many tracking cookies were used by the average website.
This is the sort of AI I want; a true personal assistant, not a bullshit generator.
That Siri went from useful to far less useful had more to do with the aim to push products at you rather than actually accomplishing the task you set for Siri. If Apple actually delivers an assistant that works locally, doesn't make me the product, and generally makes it easier to accomplish my tasks, then that's a product worth paying for.
When anyone asks "who benefits from 'AI'?" the answer is almost invariably "the people running the AI." Microsoft and OpenAI get more user data, and subscriptions. Google gets another vehicle for attention-injection. But if I run Vicuna or Alpaca (or some eventual equivalent) on my hardware, I can ensure I get what I need, and that there's much less hijacking of my intentions.
So Microsoft, if you're listening: I don't want Bing Chat search, I want Cortana Local.
There are definite frustrations, mostly around playing music. Around 5% of the time, Siri will play the wrong album or artist because the artist name sounds like some other album name, or vice versa. I wish, here, that it used my Music playback history to figure out which one I meant
Once you have the intents parsing, it should be just a matter of throwing man power at it and giving it better intents.
Yes, I have experience with building on top of such a system.
It would be easy to write Siri again and make it a hundred times better, if you could start all over and only write the core features, and not have to validate against the whole product/feature matrix.
The problem with the rewrite of course would be that you won't be able to deliver that minimal viable product any more and you will have 10 years worth of product requirements and user expectations that you MUST hit for the 1.0 release (which must be a 1.0 and not an 0.1).
I've worked on lots of "simple" and "not rocket science" systems that were 10-years old, and it is always incredibly difficult due to the state of the code, the lack of resources, and the organizational inertia.
And the AI team seemed to NOT want to be hidden in terms of discussion of industry ideas with peers at other firms.
They went back to google. Magic 8 ball time.
Who have they hired since then?
Who knows?
Who in leadership is allowing them to succeed?
Results unclear try again.
Who has a clear vision of why to build ALL ONBOARD the users device?
Results unclear try again.
This is already felt in use of Stable Diffusion, where M2 is fully capable offline.
Anything that can be done to reduce the need to “dial out” for processing protects the individual.
It erodes the ability of business and governmental organizations to use knowledge of otherwise private matters to target and influence.
The potential of moving a HQ LLM like GPT to the edge to answer everyday questions reminds me of my move from Google to DDG as my default search engine.
Except it’s even a bigger deal than that. It reduces private data exhaust from search to zero, making going to the net a backup plan instead of a necessity.
Apple delivering this on device is a major threat to OpenAI, which will have to provide some LLM model with training that Apple can’t or won’t.
Savvy users will begin to leer at having to produce queries over the wire, feeding valuable data (proven by ShareGPT)
Even then, Apple will likely chose to or be forced to open up on device AI to allow user contributed apps like LORAs which would ask the question why does OpenAI need to exist?
Also fascinating the potential to do this at the Server level for enterprise. If Apple produced a stack for enterprise training it could replace generalized data compute needs, shifting IT back to local or intranet.
That sounds absolutely horrifying if you remove the "all local" part. And that part's a pipe dream anyway. Plus, when using a model you'd basically become subservient / limited to the type of data in the model, which would necessarily abide by Apple's TOS, so a couple of hundred million people would be the Apple TOS but in human form. I don't understand why apple fanboys don't get this. Apple is pretty shoddy when privacy is concerned. Are these apple employees making these posts?
Really?! I didn't think anyone here would fall for that.
Mac Mini 12-core M2, 19-core GPU, 32GB, 10Gbit, 8TB storage? $4500
Mac Studio 20-core M1, 48-core GPU, 64GB, 10Gbit, 1TB storage is $4000. 128GB of RAM is $800 more
but either Studio RAM configuration obviously spanks the M2 mini. It's sacrificing Apple's expensive storage, but with Thunderbolt 3 it's pretty academic to find 8TB or more of NVMe storage, probably 32GB of NVMe RAID[1], for less than Apple's charge of $2200 above cost of 1TB.I spent just over $2,000.
Mac mini With the following configuration: Apple M2 Pro with 12‑core CPU, 19-core GPU, 16‑core Neural Engine 32GB unified memory 512GB SSD storage Four Thunderbolt 4 ports, HDMI port, two USB‑A ports, headphone jack 10 Gigabit Ethernet
Im satauisfied.
I'm on a newer generation chip that has a lower power draw. Meets my network speed minimum. All for the price of the entry level Studio. This box is basically an experiment to see how much processing power I need. I have a very specific project that will require the benchmarking of Apple's machine learning frameworks. I want to see how much of a machine learning load this Mini can handle. Once I have benchmarks maybe the Pro will exist and I will be in good shape to shop and understand what I'm buying.
I think a Mini of any spec is a great value. The studio has a place but I'm hoping the Pro ends up being like an old Sun E450.
This Mini experiment is to help me frame the hardware power vs. the software loads.
My second suggestion for 16-core was M2, also. $100 less with 1Gb, and with 10Gb it would be $100 more than you paid. i.e. two of the 8-core M2 Minis with 24GB RAM each would do about twice as much work as the high end Mini M2 Pro alone, sometimes less than twice the work, sometimes more. The same is true of two M1 Max Studios vs one M1 Extreme Studio for the same price. 2 less powerful machines spank one more powerful machine every single time, and one M1 Extreme Studio is definitely NOT worth two M1 Max Studios, same as one 12-core M2 Pro Mini is definitely NOT worth two 8-core M2 Minis.
Everyone is drawn to "the best," and that's where Apple fleeces and makes its money. Pretty consistently forever, the best buys from Apple are never the high end configurations. We may feel secure in what our choices were, doubling down on affirming them, but we definitely pay for it.
I think the disconnect is that you are trying to get as much processing power as possible and I'm trying to understand how much processing power currently exists.
All of them are pretty weak for local training. But having reasonably powerful inferencing hardware isn't very hard at all, UMA doesn't seem very visionary to me in an era of MMAPed AI models.
Apple Silicon AMX units provide the matrix multiplication performance of many core CPUs or faster at a fraction of the wattage. See eg.
https://explosion.ai/blog/metal-performance-shaders https://github.com/danieldk/gemm-benchmark#1-to-16-threads
Plus, the benchmark you've linked to is comparing hardware accelerated inferencing to the notoriously crippled MKL execution. A more appropriate comparison would test Apple's AMX units against the Ryzen's AVX-optimized inferencing.
For raw ML performance in a hyper-optimized system, UMA is not a big deal. For a company that needs to ship millions of units and estimate demand quarters in advance, it seems like a pretty big deal.
Apple Silicon doesn't just share address space with memory mapping, it's literally all the same RAM, and it can be allocated to CPU or GPU. If you get a 96GB M2 Mac, it can be an 8GB system with 88GB high speed GPU memory, or a 95.5GB CPU system with a tiny bit of GPU memory.
Apple's GPUs are slow today (compared to state of the art nvidia/etc), but if Apple upped the GPU horsepower, the system arch puts them far ahead of PC-based systems.
I can't believe anyone is arguing that bifurcated memory systems are no big deal. Are you like an x86 arch enthusiast? I'm sure Intel is frantically working on UMA for x86/x64, if that makes it more palatable. Though they'll need on-die GPU, which might get interesting.
As I'm working on AI stuff right now, I have to be a realist. I'm not going to go dig up my Mac Mini so my AI inferencing can run slower and take longer to set up. Nothing I do feels that much faster on my M1 Mini. It feels faster than my 2018 Macbook Pro, but so did my 2014 MBP... and my 2009 x201. Being told to install Colima for Docker with reasonable system temps was the last straw. It's just not worth the hoop-jumping, at least from where I stand.
So... when a day comes where I need UMA for something, please let me know. As is, I'm not missing out on any performance uplift though.
> I'm sure Intel is frantically working on UMA for x86/x64
Everyone has been working on it. AMD was heavily considering it in the original Ryzen spec iirc. x86 does have an impetus to put more of the system on a chip - there's no good reason for UMA to be forced on it yet. Especially at scale, the idea of consolidating address space does not work out. It works for home users, but so does PCI (as it has for the past... 2 decades).
It's just marketing. It's a cool feature (they even gave it a Proper Apple Name) but I'm not hearing anybody clamor for unified memory to hit the datacenter or upend the gaming industry. It's another T2 Security Chip feature, a nicely-worded marketing blurb they can toss in a gradient bubble for their next WWDC keynote.
Agree these are tremendously good features and having them run locally will provide the best possible experience.
Fortunately, Apple isn't an adolescent, so it doesn't suffer from FOMO.
because they find cloud based personalization so abhorrent.
I'm not even sure what "cloud based personalization" means to the user, other than "Hoover up all of your personal information."
It means having actually good ML.
I see so many posts around here saying Apple is absolutely well positioned to dominate in ML. It's just not true.
Nobody who is a top AI player wants to work at Apple where they have few if any AI products, no data, don't pay particularly well, not a big research culture, etc. etc.
The only thing they have going for them in this space is a good ARM architecture for low power matrix multiplication.
Is that a minor advantage? I would think that, the smaller the nodes get, the larger the impact of a 1nm difference. Because transistors have area, I think the math, in ≈transistor count would be 3nm:4nm = ⅓²:¼², and that’s 1,777… so a 3nm node could have 75% more transistors on a given die area than a 4nm one (roughly).
Also, most folks seem to have gone directly from 5nm to 3nm, and skipped 4nm altogether.
Is this not a battle-tank-to-battle-tank comparison?
The larger point is that Apple's lead doesn't extrapolate very far here, even with a generous comparison to a last-gen GPU. It will be great at inferencing, but so are most machines with AVX2 and 8 gigs of DRAM. If you're convinced Apple hardware is the apex of inferencing performance, you should Runpod a 40-series card and prove yourself wrong real quick. It's less than $1 and well worth the reality check.
I'm not pretending the Apple chips are the be-all-end-all of performance. They certainly have limitations and are not able to compete with proper high end chips. However I can confidently say that on mobile devices and laptops, competition is largely behind. Sure a 1000+$ standalone GPU will be faster, but it doesn't fit in my jeans. It's the same as comparing a Hasselblad camera with the iPhone 14 pro...
How can you justify your claim that they're "largely behind"? It sounds to me like the competition is neck-and-neck in the consumer market, and blowing them out at-scale. It's simply hard to entertain that argument for a platform without CUDA, much less the performance crown or performance-per-watt crown.
(Edit: In terms of privacy as there are benefits in terms of speed and offline work)
It's not like we're not already storing all of our media on the cloud (including voice), passwords and other sensitive data.
Apple has been on both sides of that coin and what is ethical isn't always clear.
> what is ethical isn't always clear
So...?
I think most users care if actually given the choice.
Even with something as simple as dictation, when iOS did it over the cloud, it was limited to 30 seconds at a time, and could have very noticeable lag.
Now that dictation is on-device, there's no time limit (you can dictate continuously) and it's very responsive. Instead of dictating short messages, you can dictate an entire journal entry.
Obviously it will vary on a feature-by-feature basis whether on-device is even possible or beneficial, but for anything you want to do in "real time" it's very much ideal to do locally.
Edit in response to your edit: nope, on privacy specifically I don't think most users care at all. I think it's all about speed and the existence of features in the first place.
I think most don’t, but they do care about latency, and that’s lower for local hardware.
Of course, it’s also higher for slower hardware, and mobile local hardware has a speed disadvantage, but even on a modern phone, local can beat in the cloud for latency.
I have to assume this is due to connectivity issues, there is no other logical reason why it would take so long to figure out what I said for so long, or not have the data on what the time is locally.
Personally, I don't trust corporations. Their motive is always money.
How do you eat?
If anyone were to write a chronological history of regulations imposed by different authorities throughoug history I think that it is a fair assumption to make that regulations related to making bread would already show up in the first chapters of the book.
And some companies, like Apple, forego short term profits from e.g. selling customer data, because it would undermine customers’ trust in them.
I would never, for instance, rely on Apple's "Translate" when Google is available.
> In a world dominated by iPhones and Samsung phones, Google isn't a contender. Since the first Pixel launched in 2016, the entire series has sold 27.6 million units, according to data by analyst firm IDC -- a number that's one-tenth of the 272 million phones Samsung shipped in 2021 alone. Apple's no slouch, having shipped 235 million phones in the same period. [1]
[1] https://www.cnet.com/tech/mobile/why-google-pixels-arent-as-...
The number of phones Google has sold is completely irrelevant to the fact that they too do local ai and have hardware on device for processing it.
How will they make money? For Apple, device purchases make local processing worth it. For Google, who distribute software to varied hardware, subscription is the only way. For reasons from updating to piracy, subscription software tends to be SaaS.
If the answer is no, then does it make sense for them to allocate those resources for such a small segment, and potentially alienate its users that choose non-Pixel devices?
Also, if the answer is no, this is where Apple would have the upper-hand, given that ALL iOS devices run on hardware created by Apple, giving some guarantees.
(I don't know this answer, it's a legitimate ask)
Part of that unique position is already being a popular product. Google adding a bunch of local ML features isn't going to move the needle for Google if people aren't buying Pixels in the first place for reasons that have nothing to do with ML.
If Google's trying to roll out local ML features but 90% of Android phones can't support them, it's not benefiting Google that much. Hence, Apple's unique position to benefit in a way that Google won't.
I've wanted to buy a Pixel for years but Google doesn't distribute it here. It's not like I'm living in some remote area, I live in Mexico, right next door.
The first couple of years I assumed Google was just testing the waters, but after so many Pixel models I suspect it's really just more of a marketing thing for Android. They don't seem to have any interest in distributing the Pixel worldwide, ramping up production, etc.
So while pixel phones may be possible, they don’t want to.
Take image processing for example. iPhones will tag faces and create theme sets all locally. Google could too, but they don’t. They send every picture to their cloud to tag and annotate.
I like this because Apple the corp doesn’t know the individuals in photos and processing happens locally.
I don't see why client-side processing mitigates the privacy concerns. That doesn't stop Apple from "locally" spying on you then later sending that data to their servers.
Also since Apple is built around selling expensive devices and services you could also see why they’d have much less incentive to spy and collect data than, say, Google or Facebook?
The cynicism of “everything is equally bad so why care” is destructive.
The only unique Apple thing here is how bad their AI products here and how behind they are in AI. This is the only thing that matters here - performance is adequate or better for the other processors out there, but you can't get anywhere without the appropriate software. Maybe they'll get smart enough to buy some AI startups/companies to get the missing talent.
(Aside, their image manipulation and Map is worse - though with Maps I dunno what's the underlying issue, and OCR was already mostly solved. I'm far from a photography expert so can't compare there).
I'll say though that no multibillion company is under existential threat. Not Apple, not Google and not even Intel. At worst they will lose a couple tens of billions and some marketshare. Even IBM still exists and took a long long time to fall to where it is still today.
They have done a ton with ML though. Some of these accessibility features, the ipad pencil, FaceID, image cataloging, live text, etc. etc. etc. showcase how Apple can not only do ML well but also make good use of them. All of it is done on device. LLMs and image generation are other examples of ML processes that Apple could include in the OS and run locally. With all of the issues surrounding LLMs and the like I am perfectly happy that Apple has been taking its time implementing them. It does feel like they could flip a switch when the time is right and that is why people say they are in a great position.
Is there any information on how they are doing with respect to local AI compared to other Smartphone SoCs? (Snapdragon, Tensor, Kirin, Helio etc)
Nothing special about Apple with regards to AI. M2 beats x86 in power efficiency, but not significantly better than other ARM processors.
Non-scientific example: for inference, whisper.cpp links with Accelerate.framework to do fast matrix multiplies. On M1, one configuration gets ~6x realtime speed, but on a very beefy AWS Gravatron processor, the same configuration only achieves 0.5x realtime, even after choosing an optimal threadcount, even linking with NEON-optimized BLAS. (Maybe I'm doing something wrong though).
My Kiwi accent needs this so bad.
Live Speech: I actually answer unknown phone numbers (usually) and would like text-to-speech on my calls because I've started to get concerned about what can be done, fraud-wise, with even small samples of my voice. So in this case, using another's voice is fine, even preferred. (Edit: I suppose I'd actually prefer a voice-changer here, which is less related to this accessibility feature. But I think Apple is unlikely to do that.)
Personal Voice: When my girlfriend texts me when I'm out with my AirPods, I think we'd both like me to hear her message in her actual voice rather than Siri's. This feature doesn't allow for that yet, but the pieces are all there.
Finally, Apple needs to detect and tell me when I'm listening to a synthetic voice, including when it's not being generated by Apple. There's fraud potential here. I'm clearly excited about this tech, but I want to know more on this front.
That's genius!
Neat idea, same with carplay, having it reliably imitate the voice of the person who sent the message would make it a lot nicer.
Though they would need to get all the TTS and intonation right first, which IME is not the case, I think having the right voice but the wrong intonation entirely would be one hell of an uncanny valley.
There's already some version of this in "Hey Siri" detection -- if I record myself with a prompt and play it back, my HomePod briefly wakes up at the "Hey Siri" but turns off mid-prompt. I guessed it was some loudspeaker detection using the microphone array, but it could be a mix of both?
And then you can use voice to text to text her back, and she can hear it in your voice! It's just like a phone call from 30 years ago, but one that requires infinitely more processing power!
There are folks who couldn’t, for a variety of reasons, do that thirty years ago. This feature is for them. The rest of us get to e.g. more naturally text a response to a call we’re listening into on a flight.
Voice memos? No. I would have to, at some point, speak it. If you’re referring to text-to-speech, there is a difference between having your speech read in a different voice and your own.
And latency!
See it as data compression on the wire.
It's worth noting this is already how iPhones work and people already love it. What I'm suggesting additionally is substituting Siri's voice for a DIFFERENT customized synthetic voice in a very specific circumstance. I'm not advocating for using synthetic voices where there currently aren't any here.
Disagree that it's the same at all. Sending discrete messages at your leisure is quite a different experience than a real-time conversation.
Other consequences of this feature:
- typos and autocorrect become even more hilarious
- you can still use her voice after you break up; heck, you probably don't ever need to date the person
- at this point: why wouldn't you pick Scarlett Johansson?
Which is potentially quite problematic as you could abuse this for all sorts of social engineering purposes.
This vector was recently highlighted as a weakness of the Australian "Voiceprint" system(1)
It's also an excellent point to make that these tools can be useful for everyone. I find myself using a number of the accessibility tools simply to speed up some of my common interactions with the phone and watch.
I've also noticed that these technologies end up in other products. For example livetext is now a standard feature on macOS/iOS yet it's an accessibility feature originating in the screenreader to deal with text flattened in images. This technology sharing also gives us a bit of a preview of what they're working on (e.g. AR)
(1) https://www.theguardian.com/technology/2023/mar/16/voice-sys...
I'm not saying Android is any better. We get a new Braille HID standard, in 2018, and the Android OS doesn't support it yet. So what does the TalkBack team have to do? Graft drivers for each Braille HID display into TalkBack, Android's screen reader like VoiceOver, with another driver coming in TalkBack 14 cause of course they can't update Android Accessibility suite like they do Google Maps, Google Drive, Google Opinion Rewards, and even Google Voice gets an update every few weeks. I mean, the accessibility suite is not a system app. If it were, it could just grab Bluetooth access and do whatever. But it's not, so it should be able to be updated far more frequently than it is. It's sad that Microsoft out of all these big companies that'll talk and talk and talk, which doesn't always include Apple, that has HID Braille support in the OS. Apple has HID Braille support too. Google doesn't, though, neither in Android or ChromeOS. They just piggy-back off of BRLTTY for all their stuff.
There's dozens of us who revel in polish, fit, and finish.
It would be nice if this was an option but I can't really fault them for not including it since it's kinda niche.
This seems to be have been fixed in 16.4, but before that I would:
1. Pull down to show Spotlight search.
2. Type “alarm” and tap “Create Alarm”.
3. Set a time and tap “Done”.
The feedback message would tell me the alarm was set to a different time, with multiple hours of difference.
This is using only first-party features to do a basic task and even that didn’t work right. I could reproduce it reliably. Luckily the message was correctly showing the wrong time that was set and I double-checked it to notice.
All the other issues I mentioned still have open feedbacks about it. All but one are regressions, the other is a security flaw which has always been there.
This sounds obvious, but imagine all the combinations of all the features that can interact, and then imagine having to test them all manually, because there's no automated model for specific Braille displays.
For decades, integration testing teams outnumber developers. It's just a hard problem, particularly at Apple's scale. It's not unlikely that there are only 10's of users experiencing a bug (though this one likely has thousands), and that it would take doubling or tripling the size of teams to find all these bugs before release.
I might use this myself even though I'm not disabled.
Also see curb cuts (1)
I do wish the “favorites” view in the phone app made the headshots big like iMessage, though.
Accessibility features can help everyone.
All of these features do seem really awesome for those that need it. Particularly the voice synthesis. I honestly just want to play with that myself and see how good it is and I am curious if they would use that tech for other things as well.
The whole new simple UI I can really see being a major point. Especially if it includes an Apple Watch component and can still be synced with one for the safety features a watch has. Particularly fall detection.
Edit:
Maybe I missed it but this brings up an interesting problem. Is the Voice Synthesis only stored on the device and never backed up or synced to a new device. Can you imagine your phone breaking or you needing to upgrade after you lost your voice (but you had previously set that up) and you can no longer use it?!?
I fully understand the privacy concerns of something like this being lost. But this isn't like FaceID that could just be easily re-created in some situations. So I really hope they thought about that.
FaceID is not backed up to iCloud. That is in part because it is a local-only feature. And it is in part because people’s faces change over time, so requiring them to re-enroll their face with each new phone ensures accuracy over time.
It may also be sensor-dependent; the model produced and stored by an iPhone 14 might not “make sense” to an iPhone 16, if the hardware is different.
I can see the privacy reasons not to have this data sync, but given the reasons you would be doing it in the first place the risk of loosing those recordings or the data for the voice would be a huge risk.
I don't doubt that Apple could sync this data, I just hope that they are. I don't see anything about that happening on this document so I worry that they won't for privacy concerns.
May 18 is Global Accessibility Awareness Day. There will be many accessibility-related announcements this week.
"For users at risk of losing their ability to speak — such as those with a recent diagnosis of ALS (amyotrophic lateral sclerosis) or other conditions that can progressively impact speaking ability — Personal Voice is a simple and secure way to create a voice that sounds like them.
"Users can create a Personal Voice by reading along with a randomized set of text prompts to record 15 minutes of audio on iPhone or iPad. This speech accessibility feature uses on-device machine learning to keep users’ information private and secure, and integrates seamlessly with Live Speech so users can speak with their Personal Voice when connecting with loved ones.
“At the end of the day, the most important thing is being able to communicate with friends and family,” said Philip Green, board member and ALS advocate at the Team Gleason nonprofit, who has experienced significant changes to his voice since receiving his ALS diagnosis in 2018. “If you can tell them you love them, in a voice that sounds like you, it makes all the difference in the world — and being able to create your synthetic voice on your iPhone in just 15 minutes is extraordinary.”
I love that Apple chooses to invest in these areas of accessibility that are extremely difficult to make sustainable businesses out of without charging users exorbitant sums.
If anyone here is interested-in or working-on open-source tech for the non speaking please reach out. I’m building tech that uses LLMs and Whisper.cpp to greatly reduce the amount of typing needed. What apple has here is great, it still requires typing in real-time to communicate. Many of the diseases that take your voice also impact fine motor control, so tools to be expressive with minimal typing are super important. Details (and link to GitHub projects) here: https://scosman.net/blog/introducing_voicebox
Nice to have something to comment on that doesn’t elicit any cynicism!
"AI is gonna destroy jobs" vs. "Secure on-device AI can improve your quality of life."
Narratives matter.
If you trust them to provide a secure model, I'd wager you haven't fully explored the world of iCloud exploitation.
In theory researches and hackers can and do watch network traffic while training an on-device model. Apple doesn't want their position of "on device AI is better for the consumer" to be tarnished, so I trust them to not screw around with this.
The screenshot seems to show that you need to speak 150 sentences.
You just need better authentication that doesn’t rely purely on voice.
"Better auth" always means more surveillance.
For example, iOS's newly announced Detection Mode features have been available since Android 6. https://support.google.com/accessibility/android/answer/9031...
iOS still doesn't have anything like Direct My Call, which is so useful that I use it as an unimpaired person. https://guidebooks.google.com/pixel/optimize-your-life/use-d...
In general, just like with applications that aren't for accessibility, it is better to use a platform that lets you customize the system to fit your specific needs. We, as technologists, should understand this better than most. iOS simply fails here.
But "We, as technologists" need to understand that that customization is a hard thing for many people. Most people, especially the people that part of what is announced today is targeted at, need something that is baked into the operating system and easier to manage.
If an accessibility feature requires someone to get the help of someone else to setup it is already a failure right out the gate. They are still reliant on someone else to help manage their phone.
Now yes there are situations where this is impossible to avoid, particularly for vision impaired people since you need to first set that up (but even that there are attempts to address this by the phone setup having those systems turned on by default).
But those are the exceptions and should not be the rule for accessibility features.
Edit:
Just to be clear there is obviously a market for highly customizable accessibility tools similar to the Xbox Adaptive Controller which would need someone else's assistance with.
But not everyone needs this level of support and where possible we should be making tools that allow someone to be fully self reliant and not rely on someone else for setup.
That is something that we technologists can fix on Android. We can't do that on iOS (without finding a vulnerability).
One of the visible impacts of MVC is that changes occur in real-time: you modify a value in a dialog, and everything affected by that value instantly updates. This is already common in Mac apps (in contrast to Windows apps, which typically want you to press “ok” or “apply”), so it wouldn’t surprise me if Apple was already using a modern MVC variant. It’s a well-known pattern.
https://developer.apple.com/library/archive/documentation/Ge...
Elsewhere, the documentation contrasts the Cocoa and Smalltalk versions of MVC where all three pieces communicate directly:
https://developer.apple.com/library/archive/documentation/Co...
Like you said, this separation means you can "drive" the same model through different UIs. That's one of the things I always thought was cool about AppleScript support -- the app exposes a different interface to the same model.
The harder problem is building accessibility API's for oblivious developers, so they can retrofit at the last minute before release (as usually happens). Apple's done a pretty good job there, harvesting the semantics of existing visual UI's to adapt for voice/hearing.
The personal voice feature is one of the best uses of AI I have read about in recent times.
Kudos to Apple!!!
It would be totally within their MO to suddenly wake up, and turn this cutting edge AI stuff, which no one is quite sure what to do with, into a killer app with super high quality. Fingers crossed.
I wish Apple would produce Starfire style demos showing off their products. In context narratives showing how real people use stuff. Covering features old and new.
https://en.m.wikipedia.org/wiki/Starfire_video_prototype
I so want the voice interface future. I became addicted to audiobooks and podcasts during the apocalypse. Hours and days outside, walking the dogs, hiking, gardening.
I made multiple attempts to adapt to Siri and Voice Activated input. Hot damn, that stuff pissed me off. Repeatedly. So I'd have to stop, take off gloves, fish out the phone, unfsck myself, reassemble, then resume my task. Again and again.
How very modal.
This mobile wearable stuff is supposed to be seamless, effortless. Not demand my full attention with every little hiccup.
So I just gave up. Now I suffer with the touch UI. And aggrevations like recording a long video of my pocket lint and invoking emergency services.
Maybe I should just get a Walkman. Transfer content to cassette tapes.
> Users can create a Personal Voice by reading along with a randomized set of text prompts to record 15 minutes of audio on iPhone or iPad.
That's a monumental software feature to deliver on a consumer device, with massive potential, but there it is, as just another dot point
It's been so helpful to be able to read a short transcript of what was just said. The live caption feature works on video calls, and while multi-speaker captioning isn't perfect, it mostly works. I also really like how Android, iOS, and Windows all seem to keep feature parity among the operating systems. I wonder how Google and Microsoft will respond to this and I wonder how Personal Voice will work for communication outside of the Apple ecosystem.
But Bravo to Apple once again for doing excellent accessibility features and continuing to improve them.
Most people can't understand why they would want a computer speaking the time, it would drive them nuts. I have ADHD and no sense of time passing. It helps me offload keeping track of time. (Which is otherwise continuously looking at a clock.)
I also turn off every auto-playing feature in any app that supports it. Some types of motion can be highly distracting. That can trigger a panic attack if I'm constantly needing to redirect my focus away from it. (If this sounds strange, it causes me to feel trapped in a tiny closet. Anxiety is a bitch.)
Google's latest video conferencing iteration lets people spam flying emoji. It's a "fun" feature that is absolute hell for me. Fortunately, there is a buried setting to remove it from my view.
It also has the temp and humidity and looks attractive and runs on usb: https://www.amazon.com/gp/product/B088LZRT94/ref=ppx_yo_dt_b...
Its a good clock.
My favorite is Spoken Content > Speak Screen: you can swipe down with two fingers from the top of the screen and it'll read the written content for you. I use it to read articles in Safari while I'm brushing my teeth or walking my dog
Hmm… telegraphing a future change?
Don't get me wrong, it's fantastic to see the tech being used this way. My Dad has lost most of his speech due to some kind of aphasia-causing condition and if this had been available just a year or two ago it could've been a big help for him; it's reassuring that others earlier in the journey will benefit from this.
I do think additionally, in the case for a situation like with your dad, Apple would need to verify that he is the same person that there's hours of audio of. That strikes me as a difficult at Apple's scale.
The tech has been around a while, but is getting better and better.
If I had a job as a voice-over artist, I'd be worried.
It triggers migraines (or vestibular migraines), tension headaches and dizziness in a minority.
See https://www.change.org/p/apple-add-accessibility-options-to-..., or FlickerSense.org, or LEDStrain.org forum.
I feel that many features currently focused on accessibility will soon be integrated into the main user interface as AI becomes a more important part of our computing experience. Talking to your phone without a wake word, hand gestures, object recognition/detection, face and head movements, etc. are all part of the future HCI. Live multimodal input will be the norm within the decade.
It's taken a while for people to get used to having a live mic waiting for wakewords, and it'll take them a while to get used to a live camera (though this is already happening with Face ID), but sooner than later, having a HAL 9000 style camera in our lives will be the norm.
When my grandfather lost his ability to speak, it was frustrating for him because he was such an intelligent, articulate person. Having a disease that renders your voice useless and then having to communicate via a small whiteboard for him was not fun. He couldn't write fast enough and often times felt like a burden because we would all be waiting to see what he wanted to say.
the Live Speech features for people with ALS will be a game changer. It saddens me this wasn't around when my grandfather passed, but am optimistic others who have this horrendous disease can utilize it to continue to communicate with their families even when ALS has taken their voice.
Only thing they need now is a simplified user interface for Apple Podcasts, 1 click button to listen to the latest episode of a certain podcast on there
“Type your password to buy this item on the App Store”. Why?
Because store items can be quite expensive (particularly the microtransactions), and it's using a stored credit card. I too get bugged by the plethora of prompts, but this is a good example of when to use one. Plus, after doing it once you can tie it back to Face ID or Touch ID and not have to enter it again for quite some time.
Zero extra security forcing me to type in his iCloud password on his device, especially given I have typed it in the passed.
It’s not financial requirements but instead layziness.
E.g. if the token can’t be refreshed.
That’s so nice. I hate when stuff moves in my vision when I’m trying to read something.
Yes, and they speak Russian not only in the Russian Federation. Almost the entire former CIS could use it, but here they cut it down to only Ukrainians .
Disabled people from Kazakhstan, Belarus and other countries somehow guilty?Plus you can have multiple to suit the context you’re in
I have several consoles open in different monitors and sometimes I accidentally run commands in some because the focus is in some other one :/
And yay to the super tacky looking deep drop shadows and rounded corners. Hate them all you want, but they're good for affordance.
It only needs that permission if I want to enable that. C'mon, Apple. You know better than this.
We live in the sci-fi future we always dreamed of. Sure, it's a dystopian future, what technology.
This is about as non-dystopian of an implementation as you can get.
It strikes me as something of a halfway point to a phone version of CarPlay. A safety feature that needs immediate investment. Car focus is overly prescriptive, and focus in general doesn't make my phone really safe in the car. And DO NOT start with me about Siri.
I could rant forever, so don't let that take away from the meat of this announcement.
If not, why did Apple think I would not want to know that bit of information.
“Apple has announced new accessibility features for its devices, including Live Speech, Personal Voice, and Point and Speak in Magnifier. These updates will be available later this year and will help people with cognitive, vision, hearing, and mobility disabilities. Live Speech will read aloud text on the screen, while Personal Voice will allow users to create a custom voice for their device. Point and Speak in Magnifier will provide a spoken description of what is being pointed at. These new features are part of Apple's ongoing commitment to making its products accessible to everyone.”
By Bard in 100 words
Yes, and they speak Russian not only in the Russian Federation. Almost the entire former CIS could use it, but here they cut it down to only Ukrainians. Disabled people from Kazakhstan, Belarus and other countries somehow guilty?
As someone with a disability, these features—even those do not cater to my disability—speak to me in a much more direct way than the typical “let’s guess what the disabled people want” bucket of accessibility features.
The icons and toolbar items do not scale up to respect my accessibility display/font settings.
even windows 10 allows me fine grained control over all UI sizing elements. but Apple apparently has other priorities.