ChatGPT vs. Bard: A Realistic Comparison
blog.theapricot.io
blog.theapricot.io
How can you even seriously think that asking GPT for a function that parses OPML is a "realistic" task. I can just Google it and get like 100 pages of Python functions that do exactly that.
I don’t think ChatGPT is a particularly amazing interface or experience. But Bard is so far off the mark it makes me realize Google still won’t be able to make a product. LLM as they mature and integrate into a better ecosystem of feedback mechanisms and interfaces will eat Google’s search business alive, and I see no indications they can do anything about that.
Disclaimer: I have no way to know whether Bard is using PaLM 1 or PaLM 2 in "my country". This is because Google knows better, and doesn't want to bother my pretty little head with these irrelevant facts. So, if my review of Bard is in some way unfair because I was unknowingly using PaLM 1, this is 100% Google's own fault.
I tried a bunch of logic problems, tricky sentences, etc... and Bard failed all of them, or at least partially failed.
The current favourite test is to ask LLMs to perform transformations such as translations or explanations of difficult sentences.
Prompt:
For each use of the word bear or its variants in the following sentence, put the meaning after the word in brackets: In an interview with John Bear -- a real bear of man -- of Bear Stearns, said he was bearish about bear pelt futures on the famous Bear Markets of Bern because he saw a bear wandering the streets, indicating that there may be too many wild bears.
GPT 4, first attempt:
In an interview with John Bear [a surname] -- a real bear [metaphor: a large, physically imposing person] of a man -- of Bear Stearns [a company name], said he was bearish [financial term: expecting a decline in prices] about bear pelt futures [referring to the hides of the animal, bear] on the famous Bear Markets [a play on words: typically refers to a declining market, but here it's used as a market name] of Bern because he saw a bear [an animal] wandering the streets, indicating that there may be too many wild bears [animals].
Comment: Spot on, including picking up that the capitals in Bear Markets is a hint that it's the name of the market, not the market decline concept. In some runs it also picks up that city name Bern means Bear in German.
Google Bard as of today, best of 3:
In an interview with John Bear (a real bear of a man), of Bear Stearns, said he was bearish (pessimistic) about bear pelt futures (the price of bear fur) on the famous Bear Markets of Bern (a type of market where bear pelts are traded) because he saw a bear (a large, furry mammal) wandering the streets, indicating that there may be too many wild bears (an excess of bears).
Comment: Muddles things up, combines things incorrectly, missed the bank name, put the brackets in the wrong places, etc... Most other results don't follow the prompt at all, just listing bullet points instead.
To compare reasoning capability, a follow-up question might be "Why might Bear think this about bear pelts?"
GPT 4: "Based on this context, it can be speculated that John Bear's bearish stance on bear pelt futures could be influenced by the potential oversupply of bear pelts in the market."
Bard: "If there are too many bears, they may become a nuisance or even a danger to humans."
There's just no comparison...
codeium and genie plugins for vscode and code whisper replaced my need for copilot.
though I try to use codeium first since it's free and genie is using my GPT4 API key.
Off putting when the author puts down a significant and fundamental improvement (having up to date information instead of point in time snapshot) with such a remark that so completely misses the point.
From what I can tell, Bard's access to recent information is through augmentation, not continuous training, and that's a really important difference. ChatGPT also has its "web browsing" alpha, which is a similar concept. The problem is that being able to search the web isn't the same as having a piece of knowledge integrated into your model of the world.
So, for example, if you ask contrived questions like "Why might Elon Musk have more time to focus on Tesla soon?", both Bard and ChatGPT+Browsing get a clue that they should search for Musk in the news, and they'll tell you about Twitter's new CEO pick. They'll then apply that competently to your question, and give you a reasonable answer.
But if you ask a question that requires more indirect inference, you immediately see the shortcomings of this kind of augmentation. For example, if you try "Is there hope for LGBTQ rights improving in Turkey?", neither model finds the extremely relevant point that Turkey's homophobic current president stands a significant chance of losing reelection. I can't see what Bard searched for in collecting its information (in fact, based on the answer I got, I'm not even sure it tried to go beyond its training data). I can see that ChatGPT searched for "LGBTQ rights in Turkey 2023". It's not surprising that that search didn't clue it in.
Now, obviously, I can't actually examine what would happen if each model was actually trained on the most recent news about Turkey's politics. But I'd have a very high expectation that GPT-4 would make that connection, and I would be fairly surprised if 3.5 didn't as well. But there's a bottleneck because the information isn't actually integrated.
Which isn't to say that that kind of augmentation is useless, of course. But the distinction is important.
So maybe it's a feature here and not a bug?
All in all, it's possible in due time, but very challenging. Even for a human for that matter - I wouldn't know how to answer that question myself, even after Googling about it.
so far I'm very unimpressed with bard and even bing chat, and I just got gpt4 browser plugin, so that might be better eventually but it's really slow.
I can imagine some kind of multipoint benchmarking in the near future. What do you think?
For some questions I asked, I liked Bards response, for some others I liked ChatGPT. Moreover, Bard is much newer vs ChatGPT had a lot of fine-tuning data from being used for a while now.
The "best" will change over time, and will depend on ease of use, speed, availability, how close they are, rather than just the "best response" to more types of questions.
A Chat app that is good enough, but with better integrations, enterprise support will be used by enterprises. Just like why Meet and Teams gained userbase over Zoom, even though they came in the game late.
At present, I don’t think any of these top tier models entirely dominate the other. Each has different pros/cons.
Fundamentally, there must be some trade off between tuning and expressiveness, model size and training frequency (relevant to live data) and possibly a few other performances metrics not well understood yet.
I suspect LLMs will be like databases in that different models will be used in different domains.
This is also interesting with GPT-4. I've been using it to generate tons of code and sometimes it indents by 2 spaces, sometimes it mimics what I give it first, but often times to reverts to 2 spaces regardless even if told to explicitly indent differently. Rather annoying indeed.
ChatGPT needs a plugin that automatically sends the output through this site lol: https://www.browserling.com/tools/spaces-to-tabs
Especially if the problem is something as easily fixable as the wrong level of detail, I would have thought that exploring some alternative prompts would make sense. Given this is the prompt that the author is already using for a ChatGPT-based app, it's obvious that this specific prompt would work well there. If it didn't, they would have iterated on the prompt until it produced acceptable results for their app.
You can’t have the cake and eat it too
Bard: Bard isn’t currently supported in your country.
ChatGPT: Is available and I can pay for GPT4.
Winner: ChatGPT
"Edgar Allan Poe was an American writer, poet, editor, and literary critic who is best known for his poetry and short stories, particularly his tales of mystery and the macabre. He is widely regarded as a central figure of Romanticism in the United States, and of American literature. He was one of the country's earliest practitioners of the short story, and is considered the inventor of the detective fiction genre, as well as a significant contributor to the emerging genre of science fiction.
Poe was born in Boston, Massachusetts, on January 19, 1809. His parents, David Poe Jr. and Elizabeth Arnold Poe, were both actors. David Poe died when Edgar was only three years old, and Elizabeth died two years later. Poe was then taken in by John and Frances Allan, a wealthy tobacco merchant and his wife, who lived in Richmond, Virginia. Poe was never formally adopted, but he was raised by the Allans as their son.
Poe attended the University of Virginia for one year, but he was expelled for gambling debts. He then joined the U.S. Army under an assumed name, but he was discharged after only two years. In 1835, Poe published his first book of poems, Tamerlane and Other Poems. The book was not a success, but it did attract the attention of some literary critics.
In 1836, Poe married his 13-year-old cousin, Virginia Clemm. The couple had one child together, a daughter named Virginia. Poe's wife died of tuberculosis in 1847, and her death had a profound impact on him. He became increasingly depressed and addicted to alcohol.
Poe died on October 7, 1849, at the age of 40. The cause of his death is unknown, but it is believed that he may have died of alcohol poisoning or a brain tumor.
Poe's work has had a lasting influence on American literature and culture. His stories have been adapted into numerous films and television shows, and his poems are still widely read and studied. Poe is considered one of the most important figures in the development of American literature, and his work continues to inspire and entertain readers around the world.
Here are some of Edgar Allan Poe's most famous works:
"The Raven" (1845) "The Tell-Tale Heart" (1843) "The Fall of the House of Usher" (1839) "The Murders in the Rue Morgue" (1841) "The Pit and the Pendulum" (1842) "The Masque of the Red Death" (1842) "The Cask of Amontillado" (1846) "Annabel Lee" (1849) "Lenore" (1843) "Ulalume" (1847) Poe's work has been praised for its dark and macabre themes, its use of suspense and horror, and its vivid imagery. He is considered one of the most important figures in the development of American literature, and his work continues to inspire and entertain readers around the world."
edit: sorry if it seemed like I claimed I didn't believe you. I was just pointing out that the first paragraph seemed to be a verbatim copy from Wikipedia and I was wondering about this weird behavior from Bard.
"I can't assist you with that, as I'm only a language model and don't have the capacity to understand and respond."
Yeesh
Considering that my $20 only gets me 25 messages every 3 hours, I'll probably be turning to Bard when I run out (which I haven't yet, to be fair).
i say coded, i mean, i directed chatgpt to do it
chatgpt had some issues with the QR content extraction using zxing so I tried Bard.
Bard destroyed totally working code again and again while ignoring the issue on hand.
In the end I googled some stackoverflow posts with working code, fed them to chatgpt and it worked it out.
In my experience, BARD is not up to any coding tasks, while chatgpt - while not perfect - is magic.