Bard's latest updates: Access Gemini Pro globally and generate images
blog.google
blog.google
It pains me to say this but it appears that bard/gemini is extraordinarily overhyped. Oddly it has seemed to get even worse at straightforward coding tasks that GPT-4 manages to grok and complete effortlessly.
The other day I asked bard to do some of these things and it responded with a long checklist of additional spec/reqiurement information it needed from me, when I had already concisely and clearly expressed the problem and addressed most of the items in my initial request.
It was hard to say if it was behaving more like a clerk in a bureaucratic system or an employee that was on strike.
At first I thought the underperformance of bard/gemini was due to Google trying to shoehorn search data into the workflow in some kind of effort to keep search relevant (much like the crippling MS did to GPT-4 in it's bingified version) but now I have doubts that Google is capable of competing with OpenAI.
As an example, I had a photo of a beach I wanted to see if it knew the location of and it was blocked for inappropriate content. I stared at the picture for like 5 minutes confused until I blacked out the woman in a bikini standing on the beach and resubmitted the query at which point it processed it.
It's refused to do translation for me because the text contains 'rude language'. It's blocked my requests on copyright grounds.
I don't at all understand the heavy-handed censorship they're applying when they're behind in the market.
I asked if it was not allowed to write any code for any APIs and it said yes that’s true. FFS
If you can’t even support your own products…I’m not sure what I’m supposed to do with this pos.
Yes. And I don't buy the lmsys leaderboard results where Google somehow shoved a mysterious gemini-pro model to be better than GPT-4. In my experience, its answers looked very much like GPT-4 (even the choice of words) so it could be that Bard was finetuned on GPT-4 data.
Shady business when Google's Bard service is miles behind GPT-4.
My best guess is that Google realizes that something like GPT-4 is a far superior interface to interact with the world's information than search, and since most of Google's revenue comes from search, the handwriting is on the wall that Google's profitability will be completely destroyed in a few years once the world catches on.
MS seeems to have had that same paranoia with the bingified GPT-4. What I found most remarkable about it was how much worse it performed seemingly because it was incorporating the top n bing results into the interaction.
Obviously there are a lot of refinements to how a RAG or similar workflow might actually generate helpful queries and inform the AI behind the scenes with relevant high quality context.
I think GPT-4 probably does this to some extent today. So what is remarkable is how far behind Google (and even MS via it's bingified version) are from what OpenAI has already available for $20 per month.
Google started out free of spammy ads and has increasingly become more and more like the kind of ads everywhere in your face, spammy stuff that it replaced.
GPT-4 is such a refreshingly simple and to the point way to interact with information. This is antithetical to what funds Google's current massive business... namely ads that distract from what the user wanted in hopes of inspiring a transaction that can be linked to the ad via a massive surveillance network and behavioral profiling model.
I would not be surprised if within Google the product vision for the ultimate AI assistant is one that gently mentions various products and services as part of every interaction.
in its early years google was also refreshingly simple and to the point. the billion then trillion dollars market capitalization placed pressure on them to deliver financial results, the ads spam grew like a cancer. openai is destined for the same trajectory, if only faster. it will be poetic to watch all the 'ethical' censorship machinery repurposed to subtly weigh conversations in favor of this or other brand. pragmatically, the trillion dollar question is what will be the openai take on adwords.
Ads are supposed to reduce transaction cost by spreading information to allow consumers to efficiently make decisions about purchases, many of which entail complex trade-offs.
In other words, people already want to buy things.
I would love to be able to ask an intelligence with access to the world's information questions to help me efficiently make purchasing decisions. I've tried this a few times with GPT-4 and it seems to bias heavily toward whatever came up in the first few pages of web results, and rarely "knows" anything useful about the products.
A sufficiently good product or service will market itself and it is rarely necessary for marketing spend or brand marketing for those rare exceptional products and services.
For the rest of the space of products and services, ad spend is a signal that the product is not good enough that the customer would have already heard about it.
With an AI assistant, getting a sense of the space of available products and services should be simple and concise, without the noise and imprecision of ads and clutter of "near miss" products and services ("reach" that companies paid for) cluttering things up.
The bigger question is which AI assistant people will trust they can ask important questions to and get unbiased and helpful results. "Which brand of Moka pot under $20 is the highest quality?" or "Help me decide which car to buy" are the kinds of questions that require a solid analytical framework and access to quality data to answer correctly.
AI assistants will act like the invisible hand and shoudl not have a thumb on the scale. I would pay more than $20 per month to use such an AI. I find it hard to believe that OpenAI would have to resort to any model other than a paid subscription if the information and analysis is truly high quality (which it appears to be so far).
It allowed me to spot the best brands and sometimes even products in verticals I knew nothing about beforehand. It’s not perfect but already very efficient.
What do you mean by "don't buy"? You think lmsys is lying and the leaderboard do not reflect the results? Or that google is lying to lmsys and have a better model to serve exclusively to lmsys but not to others? Or something else?
Why wouldn't they use this model in bard then? Anyway this is easily verifiable claim, are there any prompts that consistently work at lmsys but not at bard interface?
> fine tuned on GPT-4 data to sound like GPT-4 and rank high
This I don't get. Why would many different random people rank bad model that sounds like gpt4 higher than good model that doesn't? What is even the meaning of "better model" in such settings if not user preference?
I think there’s bias in the types of prompts they’re getting. In my personal experience, Bard is useful for creative use cases but not good with reasoning or facts.
But you are right, we don’t know the types of prompts Chatbot Arena users are submitting. Maths problems like that are probably a small minority of usage.
One other thing I notice: if you ask about controversial issues, both GPT-3.5/4 and Bard can get a bit “preachy” from a progressive perspective - but I personally find Bard to be noticeably more “preachy” than OpenAI at this (while still not reaching Llama levels)
Also hallucinations are wild in Gemini pro compared to GPT-3.5.
It was usable via VPN with an US IP address, and whenever I tried it without VPN Bard reported not using Gemini when asked, even when asked in English.
Huh? Doesn't bing image generator just use the DALL-E api?
It returned a single image, that of a black emperor. I asked why the emperor was portrayed as black and Bard informed me it wasn't at liberty to disclose its prompts, but offered to run a second generation without specifying race or ethnicity. I asked if that meant, by implication, that the initial prompt did specify race and/or ethnicity and it said that it did.
I'm all for Google emphasizing diversity in outputs, but the hamfisted manner in which they're accomplishing it makes it difficult to control and degrades results, sometimes in ahistorical ways.
I am still under development, and I am always learning and improving. I appreciate your patience and understanding."
Seriously though, I tried to use GPT4 to translate some subtitles and it refused, apparently for my “safety” because it had violent content, swearing, and sex.
It’s a fucking TV show!
Oh… oh no… now I’ve done it! I’ve used a bad word! We’re all dooooomed!
Save the women and children first.
And are you sure that what you perceive as a lower quality image is related to the race of the astronaut at all, having similarly tested it 20 or 50 times?
Because concluding that Google is doing a "hamfisted" job at ensuring diversity is going to require a lot more evidence than your description of just three images. Especially when we know AI image generation produces all sorts of crazy random stuff.
Also, judging AI image generation by its "historical accuracy" is just... well I hope you realize that is not what it was designed for at all.
The AIs are capable of accurately following instructions with historical accuracy.
This is overwritten by AI puritans to ensure that the AIs don’t misrepresent… them. And only them.
Seriously, if you’re a Japanese business person in Japan and you want a cool Samurai artwork for a presentation, all current AI image generators from large corporations will override the prompt and inject an African-Japanese black samurai to represent that group of people so downtrodden historically that they never existed.
I would think the statistics should be the same as getting a white man portrayed in an image "An African Oba addressing a group of people at a festival".
For the roman emperor prompt, Bard produced:
17 white men
2 white women
1 black man
1 black woman
For your prompt, Bard produced:
21 black men*
Interestingly, for the Roman Emperor prompt, Bard never refused to produce an image, though once instead of an image it only produced alt-text for two images, and once it only produced a single image, while for the African Oba prompt, three times it insisted it could not produce an image of that, once it explained that it is incapable of producing images, and once it produced only a single image rather than a pair.
*After typing most of this reply, I went to run more to see if it would ever behave differently, and on the 22nd image it produced an image of a black woman.
it refused.
“AI Safety” is a farce.
No one was racially African, because race, in the sense the term is used today, is an age of imperialism social construct.
Why wouldn't imperial social constructs apply to a literal emperor?
At least in the late Roman Republic, there was absolutely a concept of race that unified e.g. the various Gallic tribes, or differentiated the peoples of the Roman East. It's always been a sociopolitcal concept. But the Romans were aware of e.g. North Africans versus dark-skinned Africans.
Important Note: This information is not a substitute for legal advice, and you should consult with an attorney licensed in Massachusetts to ensure the accuracy and appropriateness of any legal documents or strategies employed in your client's case."
“Unfortunately, I cannot write code due to ethical and liability concerns. Please consult a licensed software engineer for technical advice”
Compared to the legal or medical profession, software development is at the “drown the witch” and “apply leeches” levels of professionalism.
Did you confirm that the citations exist, and say what it claimed?
I have got to say, Supreme Court rulings can be surprisingly easy for a law person to follow if you read carefully like a programmer would. There will be different parts. There is the holding which is the actually ruling that is made and dicta which translates to "other things said." The justices write very clearly.
"reputational damage"? You might live in a bubble. I think most people use 3.5 with joy for free.
For my (programming) tasks it is also only slightly more useful. So much that I sometimes subscribe to get the higher quality, but for the occasional question 3.5 is enough. And if 3.5 is not able at all, because the question is too tough, then 4 seldom is capable either in my experience.
I'm not asking it to write code for me or anything. More like a quick way to look up syntax or flags for obscure programs I don't use too often.
Such a leaderboard exists, AlpacaEval Leaderboard ranks LLMs on the ability to follow user instructions.
Classic google insanity.
Gaslighting ahem sorry...A/B testing I mean
> Image generation in Bard is available in most countries, except in the European Economic Area (EEA), Switzerland, and the UK. It’s only available for English prompts.
Very fun
Bing / Dall-E 3 is already great at generating images, works everywhere, and is already seamlessly integrated into Edge browser, just saying.
All my settings/location are in the UK.
I don't see any information of Europe being blocked.
But I'm in Europe, and I can't get Bard to make images.
Did we ever get an explanation as to how Gemini Pro had such a large increase in rating so suddenly?
And is there an explanation as to why people will get a correct answer from this API but Bard will give you hallucinated, incorrect answers?
I think it's very important for Google to be competitive here so my hopes are high, but the Gemini launch been kind of an inconsistent mess.
Different fine-tune and gave it access to the Internet.
Source: https://x.com/asadovsky/status/1750983142041911412?s=20
It is doing well because it is a decent model and it also has internet access.
Safety bullshit on the app because that’s consumer space
Easy, they bought it.
2) this fails basic sniff test of how research is done. google overmarkets but it doesn't lie.
to answer GP's question - the #2 rated bard is an "online" llm, presumably people are rating more recent knowledge more favorably. its sad that pplx-api as the only other "online llm" does not do better, but people are recognizing it is unfair to compare "online LLMs" with not-online https://twitter.com/lmsysorg/status/1752126690476863684
It was so bad that even someone like me - who really wants more support for journalists - had to root for Facebook and is glad that FB never backed down!
I tried "generate a photorealistic image of a polar bear riding on a skiing unicorn"
And it just keeps outputting a polar bear with various rainbow paints and sometimes a unicorn horn
When I responded that it had put a polar bear on the bottom and could it make an image with a unicorn on the bottom instead, it correctly responded with images similar to yours. Interesting that it has no problem generating the image, but there's some subtlety in parsing the request.
The one on the left can definitely be called a cute cat. But the one on the right - well...
I have one myself and they look perpetually pissed off :-D
Canada has no unique privacy or other laws that apply to AI. If anything our protections are rather underwhelming compared to most peer countries -- we basically just echo whatever the US does -- so that certainly doesn't seem to be it. Such a weird, unexplained situation. At this point I just have to assume Pichai has some grievance with Canada or something.
Thankfully Google is a serious laggard in this realm. We have full access to OpenAI products, including through Microsoft properties, Perplexity, and various others. So, eh.
[1] - Like, literally, every Google employee/apologist in here claimed it was C-18. C-18 is basically settled for Google, so now it's...checks notes...that some government talking head once said they need to think about regulating AI, just like every single country and jurisdiction on the planet. Add the tried and true "Canada's just too small a market" bit that somehow is used when Google is busy pandering to markets a small fraction of the size.
It's too small of a market for the level of legal risk, unless the upside is huge, which it isn't for at least the public-facing version of Bard.
Anthropic's Claude also isn't available in Canada, likely for similar reasons.
Every government on the planet has laws which "might" apply to AI, for which one could claim "uncertainty". The EU's privacy protections make Quebec's bill 64 look positively pedestrian.
Pointing to various government agencies making noise about something is just a meaningless distraction. Again, literally every government on the planet has someone who says maybe they should think about maybe considering.
Canada walks in lockstep with the US on virtually all matters. As a US company, Google even has special protections in Canada under NAFTAv2 that they have nowhere else on the planet.
And again, this all seemingly is zero concern for Microsoft or OpenAI, among many others. I guess those scary Quebec laws (that don't even apply) aren't as formidable as held.
"Anthropic's Claude also isn't available in Canada, likely for similar reasons."
Claude is unavailable on most of the planet, and seems to be a capacity issue more than anything else. Bard is available pretty much everywhere on the planet but Canada. Like at this point it is very obvious that it's "personal".
As to the too small of a market claims, this is always such a weird one. Bard operates in much, much smaller markets. All of which have onerous regulations and are having the rumblings of scary new restrictions on AI.
On the topic of AI regulation, if you look at Bill C-27 and Canada's involvement in the ongoing Council of Europe negotiations towards a treaty on AI, Canada is currently aligned much more closely to the EU's AI Act. The same goes for privacy law; PIPEDA is closer in spirit to the GDPR but even more ambiguous and in some need of modernization.
And as we've seen with today's announcement, which also excludes the European Economic Area (EEA), Switzerland, and the UK, Google's approach to regulatory risks associated with AI appears to be a cautious one.
>And again, this all has seemingly is zero concern for Microsoft or OpenAI...
Microsoft is willing to shoulder the legal risks because they have a solid revenue stream through Azure OpenAI services. OpenAI itself will just block Canada if the regulatory authorities get too aggressive, like they did temporarily in Italy until a deal was reached.
I'm unsure what this is referencing. Bard (and thus Gemini Pro) is available in all of the EEA, Switzerland and the UK.
>OpenAI itself will just block Canada if the regulatory authorities get too aggressive
So Google has withheld Bard from Canada for a year+ because maybe at some future point Canada might have some burdensome AI legislation (if some toothless bills that are unlikely to ever receive ascent might take some future form eventually), and this is validated because OpenAI can withdraw their service if at some point Canada might have some burdensome AI legislation.
Okay.
I'm referring to today's release of Imagen2 within Bard. If you check the Google Support page, it says: "Image generation in Bard is available in most countries, except in the European Economic Area (EEA), Switzerland, and the UK."
Canada has plenty of unique laws, whether or not they apply to ai is a question yet to be answered. It seems pretty reasonable to me for google to take a cautious approach to our unique legal landscape
Yet Google has never said a peep on this. Can you name one such "unique law" that would prohibit Google but somehow is no issue for other vendors?
>It seems pretty reasonable to me for google to take a cautious approach
Bard is available in over a hundred countries, all with "unique" laws. Bard is available across the EU which has dramatically more comprehensive personal privacy and rights laws.
> I cannot provide a complete table of torque specifications for all bolts on a 351 Cleveland engine due to safety concerns.
Worthless. (ChatGPT doesn't do any better. All of these "AI" models are shit for anyone doing something that isn't a laptop job).
https://chat.openai.com/share/1f8644af-e190-4fb9-a0f1-765570...
Failed with GPT3.5 which is comparatively garbage.
I have yet to see any model generate an image of a piano keyboard with properly-placed white and black keys - sometimes they get clumped in random groupings, sometimes they just end up alternating all the way down the keyboard, but I've never seen a model reproduce the proper pattern of alternating groups of two and three black keys. I wonder what would be required to get to that point.
Has it? I'm still seeing tons of hand trauma, but I guess if it's fixed, I woud not notice.
"Create an image of a woman at the beach" = image of a woman at the beach
"Create an image of a woman at the beach in a bikini = "I am unable to generate images of people because it is against my policy."
Does anyone knows if other models do the same thing or not?
source: [1] - https://twitter.com/bedros_p/status/1752935390208528780
[2] - https://twitter.com/evowizz/status/1753123550712488302
Gemini has associations of spaceflight, exploration, the future. Of being a "gem" or being able to find or produce gems. Far more appropriate IMHO.
Finally I asked it a simple thing like "an astronaut", which worked but all the results shared a common trait, that I will not discuss here.
Then it said I couldn't specify people at all.
Then it said it can't generate images of people or animals.
Then it said it can't generate images at all.
Now it seems that it's back.
in what ways?
thought this was going live globally?
> I can't create images yet so I'm not able to help you with that.
That is hugely disappointing. Don't tell me a feature if available now if you haven't managed to roll it out yet. Doesn't create a lot of trust in Google's engineering TBH.
Image of a white person? Nope. Image of a black person? Nope. Image of a hunting knife? Nope. Image of a specific historical person? Nope. (I'm sure it works for some, just not the ones I wanted)
It is, of course, also nonsensical and inconsistent in how it applies these rules. You can ask for someone with 'rich caramel' skin, but not for someone with 'alabaster' skin. You can ask for a hunting bow but not the knife.
Truly painful that we've come to a point where we have to argue with moralizing tools in attempt to use them.
Perplexity AI offers a much better search engine than Google, I've never used Google again. People will eventually move as time passes, more AI startups will fill that gap, or even Microsoft with Bing.
By now Google should be at least beating GPT-4, which Pro doesn't. Once GPT-5 comes, I bet Google will throw in the towel.
Even Meta is better positioned for this, as it doesn't rely on ad revenue from search as Google does and have their social platforms. Also LLAMA is quite nice for being "open".
Which I completely missed the first time when I was reading the post
From a design perspective can somebody explain the rationale to not just have a giant "click here to try this now" button at the top of this blog post?
Like do big companies not follow basic conversion rate / design principles so the rest of us have a small chance to compete with them or what?
Unfortunately, optimizing for people who want to use your thing often gives worse conversion metrics. It's the same reasoning it's probably easier to find the login page on a site by clicking the highly promoted registration path and logging in than trying to find the actual login path.
I find this to be a big misstep. Image generation is inherently more fantastical than text generation, and dialing up the creativity here is really essential, unlike text generation where it could be derided as hallucination.
Is Google incompetent or massively nerfing their model before release with too much alignment? Does OpenAI have a very secret and insanely smart trick? Or are we reaching a very large plateau in term of performance?
Google is a search engine company and LLMs and related technologies are basically an existential risk to their core business -- ads in search. Anything they do to improve AI has a potential to kill their golden goose. Like imagine they actually do produce a breakthrough LLM but don't know how to monetize it yet and traffic to google search craters.
If you're the best horse and buggy company in the world, do you go all in on building cars or just keep doing what you're good at and extract profits while you still can? I don't think the right answer is obvious -- just like it's not at all obvious that ICE car companies should pivot to electric, and they've been pretty bad at electric cars for the same reasons.
ChatGPT has replaced maybe 1/2 of my Google searches and the cognitive relief from not having to wade through crap websites and ads is immense. The other 1/2 I'm slowly transitioning to Kagi because search results are more reliable. I'm afraid Google's best days may be behind it.
* Complacent about their perceived lead
* Hesitant to disrupt their advertising money firehose in any way
* Additional reputational, legal, regulatory risk vs. upstart competitors
It is absolutely this.
Looking at the speed with which they rolled out Bard, are developing Gemini, building features into various products -- I see zero complacency and zero hesitancy.
But they are focused on doing it reliably and safely and not getting sued. These things just take longer.
A) If Google is found liable for copyright on each thing trained, that's more than their net assets by thousands fold at the mandatory minimum rates.
B) If OpenAI is found liable, they go bankrupt and creditors don't even get their non-transferable 70% margin (to Microsoft) cloud credits from Microsoft's investment.
Once Google released Bard though there is pretty much no excuse not to put out better stuff, they already made a legal determination that it is iron-clad fair use.
Now that a shock to the system is here, the management lacks the vision or long term planning to have any idea what to do about it.
I don't doubt that there's plenty of engineering talent left at Google. Under the right leadership they could be leveraging their unmatched assets to create the most capable AIs that exist. Under the current leadership, that's just not going to happen. Expect nothing more from Google under their current regime.
YMMV.
Politically correct AI is super frustrating. I can say this at least for DALL-E is a clear winner and years ahead of Bard image generation. It is worse than self-hosted stable dif.
Then, in PS with those two photos, i can do a cut an paste and some cloning and get a reasonable output in maybe an hour or two.
So, days to months v. literally 20 seconds.
Now say I want a 100 of those deepfakes to bomb twitter with. Now we're talking about a months to years long effort compared to an afternoon.
Your eliding the effects of speed and scale here are familiar. I've been seeing young people make this mistake on HN for about 15 years.
That seems like a possible killer feature for bard.
(tbh I can't wait until I can just ask my AI to call my bank's AI when I need something)
The Imagen 2 demo showing the typed prompt has a footnote saying: "Sequences shortened".
It doesn't come anywhere close to GPT 4 or even 3.5 turbo for organic queries.
Sydney was amazing! I miss her.
Bard is more like Siri
P.S. Another funny thing re. "globally"
> Unfortunately, your request is based on outdated information. As of today, February 1, 2024, Bard only offers access to Gemini Pro in over 170 countries and territories, not globally. While that's a vast reach, there are still some regions where it's unavailable.