Matryoshka Diffusion Models
arxiv.org
arxiv.org
I don’t follow ML so while I understand the words I don’t have enough context to know why this is good.
The models converge much more quickly during training by using this multi-resolution paradigm. There's a figure (4) in their paper showing how quickly these models converge to a good accuracy on a standard training set at 256x256 resolution, comparing them in particular to a Latent Diffusion implementation.
They showed that their baseline models were slightly better than Latent diffusion in terms of training efficiency. However, when they pretrained their models on 64x64, they converged much much faster. The Apple models converged at around 60k iterations to a better score than where Latent Diffusion arrived at in around 300k iterations.
The caveat is that for my example above, they pretrained the model for 390k iterations on 64x64 before setting it to train at 256x256, which is indeed more than Latent diffusion got the opportunity to train for - however, those 390k iterations were using something proportional to 16 times less resources.
This would be a very good paper for our friends creating Stable Diffusion to copy when creating the next versions, it would likely save them a lot of GPU time and I bet it would improve robustness for different dimension output.
* Photo enhancement and object detection
* Text identification
* FaceID
But I also echo the sentiment that they could and should be doing more.
And I only need one example to highlight their massive shortcomings...
Siri.
> gives me the current date.
“Siri, what’s 9am Pacific Time in my local time zone?”
> gives me the current time.
I could go on…
Or the classic where you ask something simple and it looks the query up on the internet verbatim instead. Or completely bungles a clearly enunciated query.
“Hey Siri, what’s 3 * 9?” “Looking up ‘treehouse sign’ on the internet”
"I got in a car and said go. It didn't go."
"Did you try putting it into gear?"
"That wasn't the promise! I already understand my horse and it goes fast enough when I say go. Why should I now have to understand a car?"
Tools are tools! AI is not different. Like any tool, it will have limits, it will get better, it will keep having other limits.
Learn the new tool if you want to, don't learn the new tool if you don't want to. But it feels disingenuous, and your valid criticisms get lost, when you claim "The tool doesn't do what I want! and stop trying to teach me how to use it!".
If a person would understand what I mean but an AI doesn't, the AI is broken.
- I can search my photos on my phone, despite not having uploaded them to the cloud (image recognition runs locally)
- FaceID
- Excellent OCR in pretty much everything
They've been cautious about shipping generative AI features, but they're absolutely leading the field in terms of building great features using edge ML running on devices.
It's really very good and proibably saves me 5 minutes on a big shop. Very understated, non-flashy ML
2. Add List
3. Set "List Type" to Grocery
I googled it the other day and setting the type was the key.
Awesome.
https://thinkml.ai/iphone-ai-artificial-intelligence-feature...
But Apple and Google are leading the cutting edge for on-device AI.
I had to turn off half of them to stop it writing random stuff for me
Yeah it did some annoying things like correct fuck to duck... but to call it non functioning is just flat out wrong. More often than it it did what I needed it to do.
Also "friendly reminder" maybe we should ask ourselves why a certain Ad company has so much data to improve their AI at faster paces, and it isn't because they have better engineers.
It also somehow misses simple corrections. Try it right now, it doesn’t know that “ennencuiate” should correct to “enunciate”. Or even “dejline” to “decline”.
What’s worst about the swap—a-roo’ing is that it will replace a common logical word in a sentence with something seldom-used.
Like, it’d correct “overwhelmingly just not how it works” to “ontologically just not how it works”.
I can't even turn it off because I use the swipe input.
It's so funny that it'll use new, wrongly guessed information to change previous context rather than using that context to correctly guess the new information.
That being said, I think they toned it down in one of the recent updates. I remember being infuriated much less by it than a year ago.