What Meta learned from Galactica, the doomed model
venturebeat.com
venturebeat.com
But it’s not like ChatGPT was better. If I recall, RLHF actually made hallucinations slightly worse. Even today OpenAI wouldn’t claim to be able to accurately summarize papers. It’s just that there was an endless supply of “write X in the style of Y” that was like catnip for journalists and kept them busy for months.
It is significantly less impressive when you ask something you don't already know and can't Google easily.
If I asked you to build a light switch I don't need to know anything about EE to see that the light turns on/off when I flip the switch.
LLMs often fail to provide functional solutions for non trivial stuff (not impressive) - and once you know more about the domain, you usually figure out that the light switch will likely burn down your house in few hours because it's overheating.
I think at some point it just becomes a preference of your interface to the data your searching, and I've been enjoying natural language
The place where I've found it really handy is if I want to make stuff up. Things like "Give me a mundane event that could happen to a 25-year-old leatherworker for a story I am writing." or "What should I name a piece of software that allows people to collaboratively draw a landscape scene?"
What's more, you absolutely cannot trust it with that kind of query.
Its insane to me that LLMs are primarily being pitched as fact givers. That is such a bad fit... They are so much better at writing poetry and fiction in chunks, or maybe analyzing stuff from the context.
I feel the "grounding in facts" was a big mistake, and to a large extent still is. The worst case of it was symbolic linguistic and knowledge graphs. Why? Because symbols don't exist, facts and concepts and words are fuzzy boundaries, their meaning defined mostly or entirely through associations with other words. We don't learn symbolically, we don't have strong "grounding in facts" - not at the language layer.
I'd rather say that language and understanding, the way we do it, is statistical in nature, and "grounding in facts" is done through feedback.
To bypass the garbage in problem, you'll need a different AI structure from LLMs almost certainly, a proper logic model.
Based on how it's reported that they work, I assume that even without garbage in their statistical reasoning of the next word would occasionally come up with bad output. E.g. they might come up with 2 + 2 = 5 because of mathematical inputs such as "1 + 2 + 2 = 5".
I'd say that LLMs are more resistant to GIGO, as long as the fraction of garbage in training data is small - it'll look like outliers to the larger model.
"Write this proof in the style of a program"
I think there is something a lot deeper behind this
[1] https://statmodeling.stat.columbia.edu/2022/11/23/bigshot-ch...
I’m actually interested in what the internal lessons for Facebook would’ve been to see their product so replaced by a similar product only two weeks later, but unless I’m misunderstanding everything, the article doesn’t really seem to touch in it. At least not beyond the fact that people still want the model, and, that it’s part of llama and Facebooks new push for more open models.
> “The gap between the expectation, and where the research was, was too big.”
> Overall, Pineau said, “If I was to do it today, we would just manage the release.”
For me, the main reasons why Galactica was more heavily criticized than ChatGPT were twofold. For one, the more scientific aspirations that this model had from the outset, proclaimed boldly in the associated paper [0] made any factual errors or problematic output far less justifiable, compared to a model that, through its branding and design, was more conversational in nature.
That, I feel, leads directly into the second important reason. The initial launch of ChatGPT, very cleverly, targeted users of all backgrounds, making it more likely that the initial interaction the majority of users had with the model was more conversational and humorous than the type of systematic picking apart the more scientifically inclined crowd tends to partake in. This led to initial coverage of ChatGPT being filled with amazement by laypeople, drowning out a lot of criticism and more measured reactions in a sea of hype, whereas Galactica was consistently bombarded with less favorable reactions.
Overall, I still feel that the approach they took in building Galactica is one that should be explored further, and I am somewhat saddened that more focused LLMs have become less favored by researchers. I remain hopeful that as we explore the limitations of universal, conversational models, there will be a resurgence of more specialized LLMs similar to Galactica.
I tried Galactica when it came out, and I have to say that subjectively at least the results looked much inferior to what the benchmarks suggest. In the paper they claim to be substantially better than GPT-3 on their larger models, while in my personal experience even some generous queries produced garbage output. I cannot remember whether the version available at the site was the largest model, however.
If I remember correctly chatGPT was grilled and prompt hacked for months, it was the Twitter obsession of the day.
I had the feeling Meta released this with great fanfare saying "do science with it" so of course there was backlash when it was shown to hallucinate.
I mean the website still sounds like it https://galactica.org/mission/
I suspect the underlying problem was the limitations being obvious to the research team and them being excited that it was -less- limited than previous attempts in the same direction and not realising how people who didn't find the limitations it still had obvious would react.
All very predictable in hindsight but because they didn't see it as a 'real' launch I don't think people with enough of an outside view were involved, and ... I think we've all been there with the "but that's obvious" problem, so (a) you're right (b) I'm still kinda sympathetic.
This is how it was advertised on Twitter by Lecun: "Type a text and http://galactica.ai will generate a paper with relevant references, formulas, and everything."
If he wanted to advertise it as experimental and make clear its limitations he could have written something like "Type a text and http://galactica.ai will generate something that looks like a paper with made up references, incorrect formulas, and anything."
You filled in a form and downloaded them…
Not really a leak in my books
This thing didn’t actually do much aside from hallucinate about the world in a scientific voice. The bot was allegedly supposed to help scientists explore connections between different complex theoretical fields, but it didn’t have a concept of what was real or practical, and the citations were all fake.
What value did they think was going to come from this aside from misinformation? At least ChatGPT could do your homework for you.
Does venturebeat realise there was GPT3 before ChatGPT?