Apple is reportedly spending ‘millions of dollars a day’ training AI
theverge.com
theverge.com
In the age of ChatGPT, Siri not only hasn’t been improved but has been getting even dumber lately with no significant announcements towards its improvements of understanding and doing more tasks in any WWDC recently. I would take what he does with a grain of salt.
I cobbled together my own smart home voice assistant on a weekend a few weeks ago, sitting on top of the OpenAI APIs (Whisper, GPT-4), and of course using porcupine for wake word detection.
It can do things I could never get the commercial products to do properly, for example I gave it a memory: When a user command comes in, I have GPT-4 evaluate whether it can be executed immediately or requires later follow-up. When a sensor event happens, the machinery re-prompts GPT-4 with the user command backlog, the sensor backlog and the current state, and it figures things out. That way, things like "Please turn of the lights after I leave the room" now work just fine, and all it takes is an afternoon of hacking and a PIR sensor on my little DIY Homebrew-lexa wood boxes. And of course it's also much better at interpreting natural language commands "in spirit" or "creatively".
I'm sure Amazon, Google & Apple have made all of these tinkering experiments, too, but deploying LLM-backed voice services to tens of millions just isn't affordable yet, especially when you factor in risk and liability.
Huge models are running on stock laptops. There is no need to send it to cloud. They had no problem e.g. sound recognition (reacting to alarms, cough, cry etc.) running only on selected devices. IPads have M1/M2 chips. And home assistant model does not need detailed understanding of neuroscience, best haskell patterns etc.
But all transformers development is pretty fresh corpo-wise. I think having a good safe dataset, which is not infringing any copyrights etc. is really hard. And they probably have to be very careful about it since it's not "only" about getting sued, but also potentially damaging partnerships they need for tv/books/music.
Btw have you tried some locally running models instead of GPT-4 for your automation? I don't want my HA touching the Internet unless necessary for 3rd party integrations but GPT-4 sets bar pretty high.
Not yet, but the desire is there, especially because the generic GPT-4 API is quite slow in responding to more complex prompts, not to mention the privacy concerns. I think the next version of my home NAS is likely to have a GPU in it and will run things like an appropriately tuned llama2 or similar to be the brain backbone of my smart home. Feels like an obvious direction for commercial NAS to go in as well.
Can't wait for a future where we can buy a Rasberry Pi 7 with a little analog compute ASIC running local inference with ease ... intelligent controllers everywhere!
You make a very valid point about economics. My naive point of view is: one would think a cash cow like Apple could afford it even to a limited extent, but then again iCloud free tier is still restricted to just 5GB so they have never been too generous with their cloud offerings.
The boards are raised off the wood surfaces by PCB spacers (I embedded M2.5 threaded sockets into the wood) and I bought speaker grilles that bulge out a little at edge of the cylinder, so that the 4-mic array would remain fully exposed also laterally. I covered the side speakers and the mic array speaker grille with acoustic textile.
The onboard code is written in very pedestrian Python and uses porcupine and the OpenAI APIs.
It roughly works like this:
1. Capture audio frames and run overlapping frames by porcupine to perform hot word detection (overlapping avoids the problem of the hotword falling inbetween frames, at a cost to latency)
2. Once the hot word has been detected, buffer all audio frames into a command buffer until silence is detected as a stop (detecting "silence" is a bit involved, taking noise levels into account, and a few other tricks, more below)
3. The command buffer is sent to Whisper for transcription
4. GPT-4 is prompted with a system message steering it's behavior, the user command transcription and a JSON print out of the state of all devices (e.g. lights and Sonos speakers) in the home, grouped by rooms
5. Following the system message, GPT-4 replies with a JSON structure of changes it would like to make to the device state, omitting unchanged bits from the original
6. Add the sensor event and memory system described above
There's a few other tricks. To improve the audio capture, I take note of spatially where the hot word is detected (i.e. which mic in the array gets the best signal) and then capture the rest & perform the silence detection with a corresponding bias.
This is actually done in a distributed fashion over the network, so if two of the AI speakers hear the same command, only one of them will end up processing it.
They end up making mainly HTTP calls to APIs that already exist around my house. I have a second RasPi in my LED shelf (another old project, https://github.com/eikehein/hyelicht/) that doubles as a Philips Hue bridge with a zigbee dongle. That's what the DIY AI speakers interact with when making changes to the lighting.
I will say: Depending on the user command and the weather in the cloud, it's pretty slow. I've tried my best to optimize the client side for perceived user latency, but there's no way around the GPT-4 API just being pretty slow, even if it's amazingly low-friction and reliable otherwise. And 3.5-turbo just doesn't cut it for what I'm trying to do.
I'd like to get all of this out of the cloud entirely. I predict the next generation of my home NAS will have a GPU in it and try to run things like fine-tuned llama2 for the home.
Also I hope you consider posting more about your home setup as I’d love to see more.
I still can't even tell it to turn the bedroom and living room lights on in one go. Bite-sized scripted chunks only. And half of the time it thinks I want to play some stupid song. I don't have my homepods for music, I even tried to remove all songs from my account but that free U2 album keeps coming back.
The only reason I use Siri is that it's the only one where I can turn off recording my voice.
But Bard is actually pretty good.
This sounds like non-news to me. Maybe a fluff piece for Apple?
I guess 'millions of dollars per day' gets more clicks than 'apple spends less than .25% of its revenue on AI'.
- size of apple
- size of apple r&d
- size of near future opportunity and value that might be gained
"Millions" just means AI is more than 3% of R&D.
AI is not a democratising force, it is a capital concentrator.
For example, imagine if self-driving was everything that was ever promised. Sales of cars without it would plummet.
Must be a fair few for the amount of money Apple makes on every new release.
“Wah Siri sucks” “wahh 5 years”
The iPhone was in development for like 8 years before it saw the light of day and everyone shat all over it and it still redefined mobile phones. Apples Silicon was in development for like 15 years and now laptop makers don’t even have a leg to stand on.
Apple isn’t going to release a half cocked ai project to the world. Apple isn’t Google who hasn’t released anything of note in like 10+ years. They aren’t Meta pouring billions into a metaverse no one wants.
Sorry what again? What is Siri then? That's even less than a half cooked product
While I agree it hasn’t improved over the years. Doesn’t change The fact it was good when it came out.
At least Google kills something when they lose interest.
I’m sure he’s a smart man, but considering where Siri is, and the fact he’s been there for 5 years, I’m not holding my breath.
That, plus the lost years of productivity during COVID -- I'd say it's okay to forgive for not turning the ship.
I also don't buy any Covid-related excuse with this one. If anything, I personally saw at least a couple companies make good use of the "never let a good crisis go to waste" mantra, using the general chaos of the early months of Covid to make some long-standing necessary-yet-painful changes - that is, changes that pretty much everyone knew were the way forward in the long term, but had risk of hits to revenue or much higher expenses in the short term. Using the pandemic as an excuse to go all-in on AI/Siri (e.g. "People are spending more time at home/on screens and we need a better voice interface") would have been the perfect approach IMO.
Apple sometimes does incremental improvements, and sometimes does an entirely new release.
Five years to improve? Sure, probably could've done a lot. Five years to start completely new? Not enough time to make it land with a bang compared to the existing Siri.
Siri is honestly so bad, I don't use it. Every time I try to schedule a meeting it tries to incorporate it in the calendar and looks for a contact. If I say "put meet with Mike on my calendar for 2 PM on Wednesday" it'll come back with "I can't find Mike on your contacts." Unless that's changed. Then when you ask it a question its 50/50 if it answers or just gives you useless "this is what I found on the web."
I've started using the ChatGPT app with Siri and it feels like how it should. Only problem is it can't do the scheduling or other useful things.
This is the standard Apple playbook. Five years is nothing for a new Apple product or foundational technology. The M1 was based on key ARM advances that Apple started building 10 years before it was introduced. Apple worked on the project that became the Vision Pro for 16 years.
So during Steve Jobs tenure? Can I conclude - it got eliminated from his biography on secrecy grounds?
It'd be more accurate to say that biographers would not have heard anything (from Steve or any other in-the-know Apple employees) about any projects that hadn't already been unveiled by Apple.
Additionally, it seems likely that biographers would've been required to not ask about possible future projects hinted at by public resources like patents: https://www.patentlyapple.com/2023/08/apple-won-a-patent-tod...
In some sense I’ve always thought Apple is focused so much on reliability and deterministic approaches that wrapping their heads around probabilistic outcomes is harder than for other companies.
Unlike the majority of so-called ChatGPT-wrapper and VC fuelled LLM companies burning millions a day on training and scrambling to compete with $0 on-device AI models.
Apple does not need to compete with OpenAI (Although Apple has the hardware and software to do so) and can do so at any time. Google is the one that needs to. So all of the "But but but Siri", "Siri is garbage", "Apple isn't doing anything with Siri", etc isn't even the point.
On-device LLMs or AI is the thing that matters to Apple and Siri will likely transition to a hybrid offline / online system eventually. Apple is in no rush to compete.
A good chunk of the S&P can’t compete with that, let alone startups.
They're making $350 billions in yearly revenues. They're not taking AI seriously.
I have no problem with searching pictures using keywords if it’s limited to the Photos app.
So why race when you are already at the finish line? This is Apple's case and anyone who owns the hardware that they are making or creating and releasing $0 free AI models out there.
This whole article and its comments is a great big storm in a tea cup.
Salary no doubt, but I’m not aware that a team is operating at a million+ a day. Red Bulls don’t cost that much, and I doubt any of them are taking yoga classes during their shift lol
Shots fired but you’re not wrong. Wild how that doesn’t get more attention.
These numbers seem right to me.