Show HN: I Remade the Fake Google Gemini Demo, Except Using GPT-4 and It's Real
sagittarius.greg.technology
sagittarius.greg.technology
It appeared to be able to wait until the user had finished the drawing, or even jumping in slightly before the drawing finished. At one point the LLM was halfway through a response and then saw the user was now colouring the duck in blue, and started talking about how the duck appearing to be blue. The LLM also appeared to know when a response wasn't needed because the user was just agreeing with the LLM.
I'm not sure how many people noticed that on a conscious level, but I positive everyone noticed it subconsciously, and felt the interaction was much more natural, and much more advanced than current LLMs.
-----------------
Checking the source code, the demo takes screenshots of the video feed every 800ms, waits until the user finishes taking and then sends the last three screenshots.
While this demo is impressive, it kind of proves just how unnatural it feels to interact with an LLM in this manner when it doesn't have continuous audio-video input. It's been technically possible to do kind of thing for a while, but there is a good reason why nobody tried to present it as a product.
The real problem is that it's simply too computationally expensive to continually feed audio and video into it one of these massive LLMs just in case it might decide to jump in.
I was wondering if you could train a lightweight monitoring model that continually watching the audio/video input and only tried to work out when the full-sized LLM might want to jump in and generate a response.
One time I was so distracted, I missed an entire paragraph someone said to me, walked to my car, drove away, and 5 minutes later processed it.
I made this demo in 2-3 hours, and I did use the "wait until the dictation results are finalized" technique which is safer (i.e. the dictation transcription is more robust) but slower.
For another demo - https://www.youtube.com/watch?v=fxS7OKh_4vc - I kept feeding the "in progress" transcription results into GPT and that was super super awesome & fast. It would just require more work to deal with all of the different timings going on (i.e. there's the speech itself from the person, the time to transcribe, sending the request to GPT, "sync'ing" it to where the person is (mentally/in their speech) at the point where GPT replies, etc.)
But yeah. Real time/continuous talk is absolutely where it's at. Should GPT be available as a websocket...?!
Basically it starts computing a response every time a word comes out of the speech recognizer, and if it is able to finish its response before it hears another word then it starts speaking. If more words come in then it stops speaking immediately; in other words, you can interrupt it. It feels so much more natural in conversation than ChatGPT's voice mode due to the low latency and continuous listening with the ability to interrupt.
There are a lot of things that need improvement. Most important is probably that the speech recognition system (Whisper) wasn't designed for real time and is not that reliable or efficient in a real time mode. I think some more tweaking could improve reliability considerably. But also very important is that it doesn't know when not to respond. It will always jump in if you stop speaking for a second, and it will always try to get the last word. A first pass at fixing that would be to fine tune a language model to predict whose turn it is to speak.
There are also a lot of things that this architecture will never be able to do. It will never be able to correct your pronunciation (e.g. for language learning), it will never be able to identify your emotions based on vocal cues or express proper emotions in its voice, it will never be able to hear the tone of a singing voice or produce singing itself. The future is in eliminating the boundaries between speech-to-text and LLM and text-to-speech, with one unified model trained end-to-end. Such a system would be able to do everything I mentioned and more, if trained on enough data. And further integrating vision will help with conversation too, allowing it to see the emotions on your face and take conversational cues from your gaze direction and hand gestures, in addition to all the other obvious things you can do with vision such as chat about something the camera can see or something displayed on your screen.
What is that type of input called in the literature and what research has been done on it? Thanks!
The real world challenge is threefold. First, null tokens would be massively over represented in training and by extent, in outputs. Second, at a computational level, outputting a continuous stream of tokens would be absurdly expensive. Third, there is not nearly as much training data of interspersed conversations as of monologues (e.g. research papers, this comment, etc.).
Also imho, I think until the context/memory problem is fully solved we won't really see the AI as having any kind of agency. But continuous, low latency interaction would certainly feel like a step towards that.
However it might seem expensive yes, but at least it only has to respond with one token.
True conversation is going to be very interesting.
What's crazy to me is that these tools are wildly impressive without the hype. As a ML researcher, there's a lot of cool things we've done but at the same time almost everything I see is vastly over hyped from papers to products. I think there's a kinda race to the bottom we've created and it's not helpful to any of us except maybe in the short term. Playing short term games isn't very smart, especially for companies like Google. Or maybe I completely misunderstand the environment we live in.
But then again, with the discussions in this thread[0] maybe there's a lot of people so ethically bankrupt that they don't even know how what they're doing is deceptive. Which is an entirely different and worse problem.
Edit: Can someone explain the downvotes? Is there a error in my response? I'm glad to learn but I'd appreciate a bit better feedback signal so that I can do so better instead of guessing.
[0] Abadie, A. (2021). Using synthetic controls: Feasibility, data requirements, and methodological aspects. Journal of Economic Literature, 59(2), 391-425.
Given the sensitivity of the market with any company perceived to be lagging behind AI developments, I wouldn't be surprised that their stock price would drop by a couple % if the demo underwhelmed.
So the whole thing is basically a ethical question of "how far can we go with polishing our demo until it becomes an unacceptable fake?" In an ideal world you'd never want to be caught in such a dilemma.
But I really don't get it either. One of the things that really sets humans apart from other animals is our capacity to perform long term planning and prediction. Why abandon that skill instead of exploiting it to its maximum potential?
There's a specific concern because of the large sentiment that the big difference between industry and academia is that in industry your products have to actually work. Ignoring the weird premise and ignoring the different TRL context, I'm not convinced that industry needs to actually make usable products. Aren't we also all complaining about how shit isn't working adequately? Google failing at seemingly simple things, and decreased quality of search. Amazon being a shitty monopoly and exploiting that to be anti-consumerism and not dealing with obvious spam and product manipulation that could be detected by a Naive Bayes filter. Or Twitter being overrun with spam bots that also could trivially be detected from a Naive Bayes filter or my block list, but blocking them actually decreases the visibility of my tweets so I'm actively encouraged by the platform to let these bots exist and follow me and like my comments.
And why would investors want this? I don't buy that there's an exclusive desire for quarterly profits and that Wallstreet does look for a diversification of their portfolios for long term blue chip stocks as well as short term gambles. That it'd require absolute insanity for a CEO or board to allow short term incentives to drive a well established company.
I really do think there's something going on but we're afraid to ask the deeper questions because we don't know the answers but I want to be encouraging that discussion even if it is thinking out loud. But maybe that insanity exists and this frustration is a result of a demonstration of it. But I don't want to believe that because I think humans are capable of so much more. That even the average person is better than an LLM but there's just issues of communication. Because even idiocracy is driven by incentive structures so I don't buy the "lol people dumb" argument, even if using better words and wrapped up in a nice bow.
It could be the principal-agent problem. The agent (employee and management) is optimizing for short-term career benefits and has no loyalty to Google's shareholders. They can quit after 3 years, so reputation damage to Google doesn't matter that much. But the shareholders want agents to optimize for longer-term things like reputation. Aligning those incentives is difficult. Shareholders try with good governance and material incentives tied to the stock price with a vesting schedule, but you're still going to get a level of disalignment.
I suppose this is where a cult-like culture of mission alignment can deliver value. If you convince/select your agents (employees) into actually believing in the mission, alignment follows from that.
"Google Gemini Outperforms Most Human Experts & GPT-4 I Artificial intelligence I Google’s DeepMind".
It's all marketing. Same reason why satya publically postes sama + others are joining a new team at MSFT to continue should the openai thing not work out.
I'm sure the marketing team can come up with good marketing that also isn't deceitful. The question is why pull a con when you already got something of value that customers would legitimately buy?
Most marketing sells the dream, not the reality. There are just many shades of grey (although 50 tends to sell well).
I said I was skeptical of the demo but, like all developments in the field, will try it out once they release it.
So now is the time to do whatever it takes to get into the conversation ... which Google successfully did I think.
My wife and I were talking about this yesterday, and I made this exact point! I told her I’m convinced Google was deceptive like this for the Wall Street crowd and normies, because to techies and researchers who actually understand AI, the extra BS is unnecessary if the technology is legitimately impressive.
The answer is always “money”. All you have to do is think “what line of thought would lead someone to believe that by lying in this manner they’ll either lose less money or make more of it?”
Yes - GPT-4V is a beast. I’d even encourage anyone who cares about vision or multi-modality to give LLaVA a serious shot (https://github.com/haotian-liu/LLaVA). I have been playing with the 7B q5_k variant last couple of days and I am seriously impressed with it. Impressed enough to build a demo app/proof-of-concept for my employer (will have to check the license first or I might only use it for the internal demo to drive a point).
llamafile --temp 0 \
--image ~/Pictures/lemurs.jpg \
-m llava-v1.5-7b-Q4_K.gguf \
--mmproj llava-v1.5-7b-mmproj-Q4_0.gguf \
--grammar 'root ::= [a-z]+ (" " [a-z]+)+' \
-p $'### User: What do you see?\n### Assistant: ' \
--silent-prompt 2>/dev/null |
sed -e's/ /_/' -e's/$/.jpg/'
Prints to standard output: a_baby_monkey_on_the_back_of_a_mother.jpg
This is something that's coming up in the next llamafile release. You have to build from source to have the ability to use grammar and --silent-prompt on a vision model right now.Weights here: https://huggingface.co/jartine/llava-v1.5-7B-GGUF/tree/main
Sauce here: https://github.com/mozilla-Ocho/llamafile
To me, you exemplify and embody the spirit of OSS, and to top that - you seem to be just an amazing human. You are an inspiration for me and many others. And even though I know I’ll never ever get close, you make me want to try. Thank you. :)
i am curious about the running costs of something like this
Congratulations, great demo! The $0.47 bill seems reasonable for an experiment, but imagine someone doing a task of this complexity as a daily job - let's say 100x times, or a little more than 4 hours - the bill would be $47/day. It feels like there's still an opportunity for a cheaper solution. Have you or someone else experimented with e.g. https://localai.io/ ?
I remember when chatgpt was released Google was saying that they had much much better models that they are not releasing because they for AI Safety. Then theu released palm and palm 2 saying that it is time to release these models to beat ChatGPT. It was not a good model.
The they hyped up Gemini, and if Gemini Ultra is the best they have, I am not convinced that they have a better model. So this is it.
So in one year, we went from Google has to have the best model, they just do not want to release to they have the infrastructure and data and the talent to make the best model. Why they really had was nothing.
It’s completely unusable for real conversation. I’m actually in a situation where I could benefit from it and was excited to use it because I remember watching the demo and how natural it looked but was never able to actually try it myself.
Now having used it, I went back and watched their original demo and I’m 100% convinced all or part of it was faked. There is just no way this thing ever worked. If they can’t manage to make conversational live translation work (which is a lot more useful than drawing a picture of a duck) I have high doubts about this new AI.
Seems like the exact same situation to me. It’s insane to me how much nerve it must take to completely fake something like this.
Honestly: it would be fun to self-host some code hooked up to a mic and speakers to let kids, or whoever, play around with GPT4. I’m thinking of doing this on my own under an agency[0] I’m starting up on the side. Seems like a no-brainer as an application.
I did add ElevenLabs support to make it a little more snazzy sounding...
So, here it is the "Compose/Trash/Recycle Sorting Hat, Built on Sagittarious" https://github.com/n8fr8/CompostSortingHatAI
You can see a realtime, unedited YouTube demo video of my kid testing it out here: https://www.youtube.com/watch?v=-9Ya5rLj64Q
They then hyped up Gemini, and if Gemini Ultra is the best they have, I am not convinced that they have a better model.
Sundar's code red was genuinely alarming because they had to dig deep to make this Gemini model work, and they still ended up with a fake video. Even if Gemini was legitimate, it did not beat GPT-4 by leaps and bounds, and now GPT-5 is on the horizon, putting them a year behind. It makes me question if they had a secret powerful model all along
On top of that the video container usually provides synchronized audio packets.
And worse: turns out GPT-4 isn't processing pixels so much as integers representing in a position in some color space like RGB!
And worse! turns out GPT-4 isn't processing integers so much as series of ones and zeroes!
Now that this is public knowledge, I'm willing to bet this was the ugly "less than candid" truth that the board sacked Sam Altman over.
Using captions to bridge this only works to a certain extent (you're giving text descriptions of what happened in the past, not what had _actually_ happened).
https://youtu.be/zut38E-BHH0?si=oj3S3qWLgvx3743I
If you need more than that, then I’ll admit to being more cynical about big tech companies than you are.