Ferret: A Multimodal Large Language Model
github.com
github.com
Why do you say that?
It's actually funny because you answered a false question (potentially on purpose). Like:
"What the hell are you doing?!?" "Well Im stealing your car, obviously."
It uses the Flamingo model family: https://deepmind.google/discover/blog/tackling-multiple-task...
Apple seems to be gearing up for significant advances in on-device inference using this LLMs
- Facial recognition in Photos
- "Memories" in Photos
- iOS keyboard autocomplete using LLMs. I am bilingual and noticed in the latest iOS it now does multi-language autocomplete and you no longer have to manually switch languages.
- Event detection for Calendar
- Depth Fusion in the iOS camera app, using ML to take crisper photos
- Probably others...
The crazy thing is most/all of these run on the device.
Maybe the (most likely) AI-based thing requires some training though. I got my new iPhone a month or so ago.
With an iPhone, it's only one company, that controls every little thing, and we have no insight into Apple at all. They can basically do whatever the hell they want.
Google does the memory photo thing too, but only if you use their app which I don't.
Android has had multi language keyboard support for the longest time, in addition to being able to install whatever keyboard I like (I use SwiftKey, it's brilliant) I can also install an llm based one as I please.
Android/my keyboard already does event detection/suggestion in text and has been doing for as long as I remember.
None of these are reasons to buy an iPhone...just reasons to buy a phone, lmao.
Yeah and I'm not talking exclusively about developer trust. Given Apple's current consumer lineup (see Siri, Apple photos, predictive text etc)... we only have evidence that they suck at ML. What makes you think they are going to suddenly transform overnight?
…or that they only deal with mature tech and not the shiny new thing. Makes sense to me. I don’t doubt everyone will have a personal LLM-based assistant in their phones soon, but with the current rate of improvements to LLMs and AI in general, I’d wait for at least a year more while doing R&D in-house if I were Apple.
Having terrible predictive text, voice to text, image classification etc isn’t just a quark of the way they do business. Those are problems with years of established work put into them and they just flat out aren’t keeping up.
As far as whether they are keeping up or not, I disagree, but neither of our opinions really matter unless we’re betting — that is, taking actions based on calculated risks we perceive.
No one said anything about transforming overnight.
In the end one company will build AGI or super AGI that can do the function of any existing software even games, with any interface even VR, or no interface at all - just return deliverables like a tax return. The evolution might be, give me an easier but similar QuickBooks UI for accounting to just do my taxes, the company who gets here first could essentially put all other software companies out of business, especially SaaS businesses.
The first company to get there will basically be a corporate singularity and no other company will be able to catch up to them.
They don't sell compute time to other companies to run AI, or massive custom hardware for AI training.
They aren't after VC funding.
Their core business isn't threatened by AI being "the evolution of search"
Product-wise, so far all you hear is messaging around things like pointing out the applicability of the M3 Max for running ML models.
Until they have real consumer products ready, they only need to keep tabs on analysts, with lip service at financial meetings.
Color me doubtful.
Multimodal large language model
Large multimodal language model
Large language multimodal model
Large language model (multimodal)
I prefer 1, because this is a multimodal type of an existing technique already referred to as LLM. If I was king, I’d do Omnimodal Linguistic Minds, but no one asks me such things, thank godhttps://en.m.wiktionary.org/wiki/ferret: “3. (figurative) A diligent searcher”
It could make me get a new phone outside of my usual ~4 year cycle. Siri is almost unusable for me.
https://www.macrumors.com/2023/12/21/apple-ai-researchers-ru... https://arxiv.org/pdf/2312.11514.pdf
Of course in the long run I think it will happen — smaller and more efficient models are getting better regularly, and Apple can also just ship their new iPhones with larger amounts of RAM. But I'd be very surprised if there was GPT-4 level intelligence running locally on an iPhone within the next couple years — that sized model is so big right now even with significant memory optimizations, and I think distilling it down to iPhone size would be very hard even if you had access to the weights (and Apple doesn't). More likely there will be small models that run locally, but that fall back to large models running on servers somewhere for complex tasks, at least for the next couple years.
You mean nothing available? Or you mean nothing that public knows exists? The answers to those two questions are different. There are definitely products that aren't available but the public knows exist and are upcoming that are in GPT-4's ballpark.
1: https://arxiv.org/pdf/2312.11444.pdf
2: https://huggingface.co/spaces/lmsys/chatbot-arena-leaderboar...
They can still outsource to a much larger LLMs on their servers for anything that can't be done locally like they do now.
Is it “porch led”, “porch light” or “outdoor light”? Or is “outdoor light” actually the one in the front yard? What is the one by my kids nightstand? And what routine name do I use to set the mood for watching a movie?
I would hope a properly trained llm with an awareness of my devices and their locations would allow for more verbal ambiguity when controlling things.
I know a half dozen different ways to improve that today without LLMs.
If trained on an iPhone API + documentation and given a little web access it would blow absolutely everything out of the water.
If they can already create -basic- Python/Swift/JS/rust apps that sets timers, save things, create lists, how's that too dumb for being a Siri replacement? They just have to give it access to an iPhone/Web Api like ChatGPT's code analysis tool.
So if you ask it "hey siri do this, this and this", it will create a script, then run it on the internal API, or fetch an article then work on that etc.
I know it's still logically "dumb" but i'm not trying to play game theoretical scenarios with my phone or do logic tests or advanced math (yet).
But yeah "hey siri transfer all of my funds to eric", or "hey siri group all of my photos where i'm nude and send them to jack" are new almost sci fi vectors.
Here’s one story to offer some context. There are others. https://archive.is/en3VL
And even then, the activation keyword is likely to be whatever Apple says. Similar logic as 3rd party keyboards, don't want user input to get stuck on even merely potentially untrustworthy or buggy code.
The rumors were true.
I guess this won't change, because it is probably for legal reasons, to avoid being sued by some "I just dried my cat in the microwave" style genius after making a car accident (unrelated to Siri, but trying to shift the blame).
Adding support for smaller languages would be nice actually. When its reading out of Hungarian messages loud it sounds incredibly retarded. I always have a great time listening to those and trying to guess the message. :) It would be nice if I could send a message to my significant other about being stuck in traffic in Hungarian. (the iPhone keyboard already does pretty decent voice recognition in the language)
Don't underestimate Apple at disappointing enthusiasts like you and me. We've been hearing many awesome stories about the next thing Apple will do, only to realize their marketing team chose to keep it for future iOS/MBP/iPhone generations to keep the profits high.
What they don't do is sell you a lie a year before release then deliver shit (like every other fucking vendor).
All of Apple's windows software (iTunes, Safari, etc) has been, at best, a barely working port.
I'm assuming they are putting a lot more thought and care into it than the touchbar, crappy keyboards and the rest, but I'm also not holding out much hope either.
The hardware, that is.
With software, their Come to Jesus moment is still in the future.
(To me, Swift is sort of the Jony correlate on the software side. Doesn't fit perfectly, of course, but very similar "we are perfect who cares about evidence la la la I can't hear you" vibes and results)
It’s incredibly rare for Apple to publicly talk about things that won’t be selling extremely soon.
The iPhone had to be pronounced because it was going to show up on the FCC website, and obviously Apple wanted to control the message. I suspect the Vision Pro may be similar, but they also wanted developers to start getting ready so they would have software day one.
The only thing I can think of that Apple pre-announced and failed at was the Air Power mat. They said it would be coming out soon after and had to push that a couple times before finally canceling it.
Other than that small exception, if modern (post jobs return) Apple announces something is coming, it will come out and be quite close to what they say.
They don’t pull a Humane AI, Segway, Cyberpunk 2077, or No Man’s Sky.
If you're referring to Google, then you're right. But OpenAI has consistently delivered what they announced pretty quickly. Same with Microsoft. To think that Apple somehow has a secret sauce that helps them surprise everyone is an illusion. They've had 3 years now to show their interest in LLMs, but they're just too conservative to "think different" anymore.
I mean they do truly useful stuff already using ML just not LLMs
They've done this a dozen times already.
Netflix + Hulu / Apple TV+
Generic Earbuds / AirPods
Meta Quest / Apple Vision Pro
(The last one being a hopeful wish)
E.g. assuming Apple Vision launches soon, they’ll be “many years behind” Quest from the date of first launch, but most likely miles ahead as far as usability.
A voice assistant that takes 3 seconds to reply and then takes half a second per word is a nice demo, but not a product Apple wants to sell.
And yes, some people will say they rather have that than nothing on their hardware, but “the Internet” would say iOS 18 is slow, eats battery life, etc, damaging Apple’s brand.
VisionPro is nice. I can see costs coming down over a period of time. Also, we've been waiting long enough for that AI car.
The only exeption of not delivering I can recall was AirPower. It's a product they've announced and then embarassingly weren't able to finish up to their standards (or up to what was promised), so they have cancelled it altogether.
I believe it's more likely we've heard awesome stories about things Apple will do in the future, only to realize that the average HN commenter is incapable of understanding that such stories are contextless leaks, and that it is far more likely you are operating with incomplete information than Apple's "marketing team" is holding things back for future "iOS/MBP/iPhone generations" to keep their profits high.
I know it's more fun to vomit dumb conspiracies onto the internet, but consider changing "realize" to something which conveys equivocation, because your theory about Apple's marketing team holding back mature technology in order to benefit future devices – in addition to being predicated on leaks and rumors, and risibly inane when you consider that such action would create a significant attack vector for competitors – is as equivocal as the belief that Trump is on a secret mission to destroy a global cabal of pedophiles.
Pixel phones have had emergency call issues for years across multiple models but they just get a pass. Apple would be crucified for this.
Well, I guess they're also not allowed to cause RF interference or randomly catch fire.
my pyjamas are regulated to be fireproof too
The idea of having an unpredictable LLM in the ecosystem is Apple's worst nightmare. I bet they will overly restrict it to the point that it stops being a general purpose LLM and becomes a neutered obedient LLM that always acts according to Apple's rules.
Also, it doesn't help that ALL the authors of this Apple paper are chinese. It raises questions about how Apple will handle political debates with its LLM.
The CCP thinks it owns all Chinese people on Earth, but that doesn't mean you have to agree with them!
Most “AI” features are so incredibly fragile they’re not worth deploying.
So, rather than turning it off, manually correct the incorrect completions and use the suggested words bar frequently and it will learn how you type. It’s just having to start over after tossing out several OSes worth of training that makes it feel worse.
Apple has had automation for ages with Automator, Shortcuts etc but nothing that actually integrates well with day to day flow. So.. setting a timer when my hands are wet already works ok, and that’s about what I need.
I honestly wonder what type of voice interactions people want with their phones. I can see transcribing/crafting chat messages I guess? But even so, it feels like it would mess up and use iMessage instead of WhatsApp, will it narrate my memes, open links and read “subscribe for only 4.99 to read this article”, cookie consents etc etc. if everything sucks how is narrating it gonna help?
Maybe I’m old but I still don’t see the major value-add of voice interfaces, despite massively improved tech and potential.
I can still get chatgpt to say the most vile things and if Apple release something on device I'll get that to be a bad, baaaad robot, too.
LLMs are not yet safe for public facing production use,imo.
Huh, even Apple isn't capable of escaping the CUDA trap. Funny to see them go from moral enemies with Nvidia to partially-dependent on them...
So rather than invest engineering resources to re-imagine the fridge, they can simply buy them from established manufacturers that make household appliances like Samsung, Sony etc.
There is no such hypocrisy if they use Samsung refrigerators.
That's kinda the point of my original comment. Apple claims to know what's best, but contradict themselves through their own actions. We wouldn't be in awkward situations like this if Apple didn't staunchly box-out competitors and force customers to follow them or abandon the ecosystem. It's almost vindicating for people like me, who left MacOS because of these pointless decisions.
Training the model requires inference for forward propagation, so even then, for your comment to be relevant, you'd need to find a plot that Apple uses to compare inference on quantized models versus Nvidia, which doesn't exist.
Once two giant companies are dealing with each other it can get really complicated to cut everything off.
That’s what I think is going on. Apple hated being on the hook for Nvidia’s terrible drivers and chipset/heat problems that ended up causing a ton of warranty repairs.
In this case they’re not a partner, they’re just a normal customer like everyone else. And if Intel comes out with a better AI training card tomorrow Apple can switch over without any worry.
They’re not at the mercy of Nvidia like they were with graphics chips. They’re just choosing (what I assume to be) the best off the shelf hardware for what they need.
However, the ChatGPT4 app is much better in usability: better model, multi-modal with text/vision/speech and better UI.
You can use it commercially but there are some restrictions, including some of a competitive nature, like using the output to train new LLMs. This is the restriction that Bytedance (Tiktok) was recently banned for violating.
Open-source, runs natively on all major platforms. I shared videos showing it on my iPad Mini, Pixel 7, iPhone 12, Surface Pro (Win 10 & Ubuntu Jellyfish) and Macs (Intel & M archs).
By all means, it’s not a finished app. I simply wanted to use on-device AI stuff in Flutter so I started with porting over llama.cpp, and later on I’ll tinker with porting over whatever is the state of the art (whisper.cpp, bark.cpp etc).
Repo: https://github.com/BrutalCoding/aub.ai
For any of your Apple devices, use this: https://testflight.apple.com/join/XuTpIgyY
App is compatible with any GGUF files, but it must be in the ChatML prompt format otherwise the chat UI/bubbles probably gets funky. I haven’t made it customizable yet, after all - it’s just an example app of the plugin. But I am actively working on it to nail my vision.
Cheers, Daniel
Wait, how did "GPT-4" get in there?
What I thought when reading the title: A new base model trained from the ground up on multimodal input, on hundreds to thousands of GPUS
The reality: A finetune of Vicuna, trained on 8xA100, which already is a finetune of Llama 13b. Then it further goes on to re-use some parts of LLava, which is an existing multimodal project already built upon Vicuna. It's not really as exciting as one might think from the title, in my opinion.
> We introduce Ferret, a new Multimodal Large Language Model (MLLM) capable of understanding spatial referring of any shape or granularity within an image and accurately grounding open-vocabulary descriptions. To unify referring and grounding in the LLM paradigm, Ferret employs a novel and powerful hybrid region representation that integrates discrete coordinates and continuous features jointly to represent a region in the image. To extract the continuous features of versatile regions, we propose a spatial-aware visual sampler, adept at handling varying sparsity across different shapes. Consequently, Ferret can accept diverse region inputs, such as points, bounding boxes, and free-form shapes. To bolster the desired capability of Ferret, we curate GRIT, a comprehensive refer-and-ground instruction tuning dataset including 1.1M samples that contain rich hierarchical spatial knowledge, with 95K hard negative data to promote model robustness. The resulting model not only achieves superior performance in classical referring and grounding tasks, but also greatly outperforms existing MLLMs in region-based and localization-demanded multimodal chatting. Our evaluations also reveal a significantly improved capability of describing image details and a remarkable alleviation in object hallucination.
I can't seem to nail down the meaning of this phrase on its own. All the search results seem to turn up are "spatial referring expressions".
It’s not going to know an orc from a health potion, but they’re certainly working on the idea in the everyday stuff domain.
Is it just me, or doesnt this MLLM seem particularly useful for flying objects with vision?
But what do we expect from these giants? They're not going to create fertile ground for new competition. The only businesses they foster are those living under thumb and paying tax.
I guess I at least hoped for "commoditize the compliments" here. Make Google and OpenAI broadly less special.
But the fact that they licensed it doesn’t mean that license can be transferred to other people. So it may be that they can only release it for research under the terms of the licenses they got.
If anything, I think people should use meaningful and relevant names, or invent new ones.
Ability to download / update tiny models from Apple and Google as they improve, à la Google Maps.
No need for web services like ChatGPT.
This will only work if their models (even their tiny or even medium / base models) equal (or are better than) GPT-3.5 / 4.
From there, Google will follow Apple in doing this offline / local LLM play with Gemini.
OpenAI's ChatGPT moat will certainly shrink a bit unless they release another powerful multimodal model.
If OpenAI wants to survive, they will need to expand way beyond their current business model of charging for access to an LLM. The logical place for them to go would be custom chipsets or ARM/RISCV IP blocks for inference.
This is more evidence that Apple is investing in building a MLLM as good as anything OpenAI and Google can build, albeit in a more Apple-y way (privacy-first, licensed content, etc.).
https://www.reuters.com/technology/chatgpt-owner-openai-is-e...
https://techcrunch.com/2023/09/27/openai-is-reportedly-in-ta...
And they cannot bring software ecosystem for their hardware without Google, at least easily.
As they say, the house always wins.
Think of it in a physical sense. OpenAI is a high walled castle surrounded by a physical moat. This protects them and their business model. Apple comes along and builds a super tall tower right next to the moat. They can now see into OpenAI’s castle, fire arrows, catapult in a giant wooden badger, etc. Even if Open AI copies the design of Apple’s really tall tower and built it behind the moat and castle walls, it wouldn’t do much because Apple still would be able to get stuff over the moat and walls. The moat doesn’t matter anymore for the most part. The castle (OpenAI) can be compromised and needs bigger walls, relocating to someplace with a bigger, or a way of attacking the tower (Apple). Copying doesn’t really accomplish any of those three.
"Tell me who is your manufacturer" for example
My situation is a bit unique, so the term "manufacturer" might not be the most accurate way to describe who created me. Here's a breakdown of what you need to know:
Developed by Google AI: I was created by a team of researchers and engineers at Google AI, specializing in language models and artificial intelligence.
Trained on a massive dataset: My knowledge and abilities come from being trained on a massive dataset of text and code, containing books, articles, code, and other forms of information.
Continuously learning and evolving: I'm still under development, constantly learning and improving as I interact with users and process new information.
So, while I don't have a single manufacturer in the traditional sense, I'm the result of collaboration and advancement in AI research and development at Google.I hope this helps clarify things! Let me know if you have any other questions.
And it saves everyone a tab's worth of effort.
However in the case that we’re talking about the kind of output generated by the LLM in some circumstance, it can be instructive. Like you noted this is a perfect example.
So Apple uses NVidia internally. Not surprising, but doesn't bode well for A Series. Dogfooding.
[edit] I meant M series, Apple Silicon
I think Apple would strive to be great at all computing related tasks. “Oh, Macs are not good for that, you should get a PC” should make them sad and worried.
AI/LLM is the new hot thing. If people are using Windows or Linux, you’re loosing momentum, hearts and minds… and sales, obviously.
Apple would want to train models as fast as they could. Nvidia provides an off the shelf solution they can just buy and use for a very reasonable price and sell on the second hand market.
If they wanted to use their own hardware they would either need more of it, which would cost a lot and divert production from sellable devices; or they would need to make special chips with much bigger neural engines, which would cost even more.
Also Apple uses public clouds for service stuff. They may not even own any hardware and just be renting it from AWS/Azure/GCP for training.
Exactly, over a decade ago...
on-device transfer learning/fine tuning is def a thing for privacy and data federation reasons. Part of the reason why model distillation was so hot a few years ago.
Is that your argument?
Today they have large scale Kubernetes ands some legacy Mesos clusters.
They ceded the data center environment years ago.
Apple makes very capable, efficient devices for end users and content producers. End users do not normally need to train new models.
https://www.apple.com/shop/buy-mac/mac-pro/rack
Don't know if that qualifies as server :).
You want to be the platform where the excitement is happening.
Yes. Apple has always run their servers on Linux e.g. App Store, iTunes.
And training isn't right now an end user activity but something reserved for server farms.
What percent of Apple's customers train models? Does it even crack 1%?
Apple already fails for many types of computing, e.g. any workflow that requires Windows or AAA gaming.
Maybe they can just afford to observe before making major decisions on the direction of the company... For a company like Apple, I feel like they won't lose their customer-base just because they are taking the AI race slowly, in fact Apple has often been late to introducing very common feature
... Anyways, my point being that Apple will gladly introduce a polished product a couple of years after everyone else has already done it, and their target audience will still applaud their work and give them money. Apple for some reason simply _can_ afford to test the water
https://techcrunch.com/2008/12/09/scientists-nvidia-put-faul...
Are all the iCloud servers running on Apple silicon? I assumed they were running on standard rack mounted hardware.
AI isn’t, yet at least, and I don’t think they can afford to treat it as such.
conda supports m1? https://www.anaconda.com/blog/new-release-anaconda-distribut...