That's your response? Ouch.
That's your response? Ouch.
Google essentially claimed a novel approach of native multi-modal LLM unlike OpenAI non-native approach and doing so according to them has the potential to further improve LLM the state-of-the-art.
They have also backup their claims in a paper for the world to see and the results for ultra version of the Gemini are encouraging, only losing in the sentence completion dataset to ChatGPT-4. Remember the new Gemini native multi-modal has just started and it has reached version 1.0. Imagine if it is in version 4 as ChatGPT is now. Competition is always good, does not matter if it is desperate or not, because at the end the users win.
Don't buy into marketing. If it's not in your own hands to judge for yourself, then it might as well be literally science fiction.
I do agree with you that competition is good and when massive companies compete it's us who win!
[1] Google’s Bard chatbot is getting way better thanks to Gemini:
https://www.theverge.com/2023/12/6/23989744/google-bard-gemi...
There is nothing in any of Google's claims that preclude the architecture being the same kind of composite system. Maybe with some additional blending in of multimodal training earlier in the process than has been published so far. And perhaps also unlike GPT-4V, they might have aligned a pretrained audio model to eliminate the need for a separate speech recognition layer and possibly solving for multi-speaker recognition by voice characteristics, but they didn't even demo that... Even this would not be groundbreaking though. ImageBind from Meta demonstrated the capacity align an audio model with an LLM in the same way images models have been aligned with LLMs. I would perhaps even argue that Google skipping the natural language intermediate step between LLM output and image generation is actually in support of the position that they may be using projection layers to create interfaces between these modalities. However, this direct image generation projection example was also a capability published by Meta with ImageBind.
What seems more likely, and not entirely unimpressive, is that they refined those existing techniques for building composite multimodal systems and created something that they plan to launch soon. However, they still have crucially not actually launched it here. Which puts them in a similar position to when GPT-4 was first announced with vision capabilities, but then did not offer them as a service for quite an extended time. Google has yet to ship it, and as a result fails to back up any of their interesting claims with evidence.
Most of Google's demos here are possible with a clever interface layer to GPT-4V + Whisper today. And while the demos 'feel' more natural, there is no claim being made that they are real-time demos, so we don't know how much practical improvement in the interface and user experience would actually be possible in their product when compared to what is possible with clever combinations of GPT-4V + Whisper today.
Perhaps for audio and video is by directly integrating the spoken sound (audio mode -> LLM) rather than translating the sound to text and feeding the text to LLM (audio mode -> text mode -> LLM).
But to be honest I'm guessing here perhaps LLM experts (or LLM itself since they claimed comparable capability of human experts) can verify if this is truly what they meant by native multi-modal LLM.
Also I guess I don’t see it as critical that it’s a big leap. It’s more like “That’s a nice model you came up with, you must have worked real hard on it. Oh look, my team can do that too.”
Good for recruiting too. You can work on world class AI at an org that is stable and reliable.
I think it's app only though
Though now that I am reading the Gemini technical report, it can only receive audio as input, it can’t produce audio as output.
Still based on quickly glancing at their technical report it seems Gemini might have superior audio input capabilities. I am not sure of this though now that I think about it.
Multimodal would be watching YouTube without captions and asking “how did a certain character know it was raining outside?” Based on rain sound but no image of rain
From https://bard.google.com/updates:
> Expanding Bard’s understanding of YouTube videos
> What: We're taking the first steps in Bard's ability to understand YouTube videos. For example, if you’re looking for videos on how to make olive oil cake, you can now also ask how many eggs the recipe in the first video requires.
> Why: We’ve heard you want deeper engagement with YouTube videos. So we’re expanding the YouTube Extension to understand some video content so you can have a richer conversation with Bard about it.
Just don’t speak to xooglers about it. ;)
Time to press some keys on my keyboar-
> Just don’t speak to xooglers about it. ;)
Oh shit, nevermind, you get it.
You know those stats they're quoting for beating GPT-4 and humans? (both are barely beaten)
They're doing K = 32 chain of thought. That means running an _entire self-talk conversation 32 times_.
Source: https://storage.googleapis.com/deepmind-media/gemini/gemini_..., section 5.1.1 paragraph 2
Google being more open here about what they do is in their favor.
Extended musings on it, please ignore unless curious about evolution patterns of memes:
Funnily enough, it's gotten _easier_ to talk about over time -- i.e. on day 1 you can't criticize it because it's "just a rumor, how do you know?" -- on day 100 it's even worse because that effect hasn't subsided much, and it spread like wildfire.
On day 270, the same thing that gave it genetic fitness, the alluring simplicity of "ah yes, there's 8x going on", has become the core and only feature of the Nth round of the telephone game. There's no more big expert-sounding words around it that make it seem plausible.
Silicon Valley hates Canada.
It screams desperation to be seen as ahead of OpenAI.
Sounds like it's you that needs to calm down a bit. God forbid we get some competition.
It would be funny if it only happened 10 or 20 times.
Im sure they will deliver a great api for this ai then change it in a way that breaks everything.
You will fix yours I will delete mine. I will feel dumb. You will improve your life again and again basically 20+ times. Enjoy!