ChatGPT is not all you need. A SOTA Review of large Generative AI models
arxiv.org
arxiv.org
Also, what is "SOTA" for review? There isn't exactly a benchmark to compare...
In terms of comprehensiveness they don't mention PaLM + variants, which probably should be mentioned as it is currently the largest LLM with SOTA on several benchmarks (e.g. MedQA-USMLE).
In terms of correctness, I admittedly skipped to the sections I'm familiar with (LLMs) but I don't understand why they are distinguishing 'text-science' from 'text-text', they're both text-text and there is no reason why you can't, for example, adapt GPT3.5 to a scientific domain domain (some people even argue this is a better approach). A lot of powerful language models in the biomedical domain were initialized from general language models and use out-of-domain tokenizers/vocabularies (e.g. BioBERT).
The authors also make this statement regarding Galactica:
"The main advantage of [Galactica] is the ability to train on it for multiple epochs without overfitting"
This is not a unique feature of Galactica and has been done before. You're allowed to train LLMs for more than 1 epoch and in fact it can be very beneficial (see BioBERT as an example of increasing training length).
People GENERALLY don't do this because the corpus used during self-supervised training is filled with garbage/noise, so the model starts to fit to that instead of what you desire. There is nothing special about Galactica's architecture that specifically allows/encourages longer training cycles but rather they curated the dataset to minimize garbage. As another example, my research involves radiology NLP and when doing domain adaptive pretraining on a highly curated dataset we have been going up to 8 epochs without overfitting.
I believe it's a reference to 'attention is all you need,' which is a very famous paper now. I'm guessing you probably already know that -- did you cringe with it? It was a landmark paper so at least it's worth hyperbole in the title. Maybe I'm missing that its usage started before that.
Reminds me of my ML prof who would complain about people using the word 'optimum' or 'optimized' in titles of papers... one could almost always optimize more, and with respect to what is not specified by the title.
Not sure that every other AI-related paper needs to copy this format, this trend is the academic equivalent of click-baiting (seemingly trying to associate with the Vaswani et al paper) and in my anecdotal experience the usage of this play on words seems inversely correlated with the paper's quality.
Dr. Strangelove Titles - "<product>: or how I learned to stop <doing something> and love <idea>"
Dairy Farmers of America reference - "Got <Product>?"
Breakin' 2 reference - "<Product> 2, Electric Boogaloo"
Honorable Mention, titles that end in phrases like "And That's a Good Thing!", "Here's What That Means", "This is Why That Matters"
Can you elaborate? (I'm not a native speaker)
Afaik it wasn't used prior to it. While attention (and residual connections) were used with recurrent / auto regressive models, the paper showed that an encoder-decoder (non auto regressive and auto-regressive) architecture with just attention (and residuals) is sufficient to achieve great results given that you provide positional embeddings.
So prior to transformers it was RNN + Attention (+ Residuals), but the paper claimed that Attention (+Residuals+Pos Embeddings) is all you need.
"ChatGPT is not all you need. A State of the Art Review of large Generative AI models"
Verse 1: Sell the kids for food Weather changes moods Spring is here again Reproductive glands Verse 2: A country battle song Multiply, exciting people Come on, join the party Come on, everybody Chorus: In bloom In bloom In bloom In bloom Verse 3: Subhumanity is fun You have one, you have none A soap impression of his wife Which he ate and donated to the National Trust Verse 4: I'm not like them But I can pretend The sun is gone But I have a light Chorus: In bloom In bloom In bloom In bloom
Similarly, don't judge it based on its ability to solve math problems.
When you use it for the many tasks it IS suitable for it's really, really impressive.
I'm curious how this is going to play out over time. They way I interact with ChatGPT is very very different from how I interact with Google. When I try to use them the same way one of them fails in frustrating ways.
> When you use it for the many tasks it IS suitable for it's really, really impressive.
I completely agree with this. For me I've found ChatGPT to be very useful in helping to generate ideas, learn about things, or explore topics that I don't fully understand. Basically if I'm curious about something, then ChatGPT is really useful. The more I use it, the more I find that I'm poking at responses with follow up questions. It feels much more like a conversation and my interactions/approach is changing as a result. I'm now finding my approach with ChatGPT to start broad ("What's the difference between ____ and ____?"), followed by more questions to dive deeper ("Can you tell me more about ____ and provide some examples?"). It does have it's limits, but I find the way of interacting and teasing out the details that I'm after to be far more interesting and useful. And the more I use ChatGPT, the less I want to use Google. At this point, Google is mostly just simple searches only like "which service is streaming ____?" or "restaurants in my area".
Question: Yesterday evening a tree had 5 apples, a bird ate an apple from the tree at midday yesterday. No other apples were eaten from the tree. How many apples were on the tree yesterday morning?
Response: There were 5 apples on the tree yesterday evening. If a bird ate one apple from the tree at midday yesterday, then there were 5-1=4 apples on the tree yesterday morning.
Roleplay exercises - practice difficult management conversations with an employee, for example.
Brainstorming ideas.
Explaining well known concepts - especially if you ask it to invent analogies or mnemonics.
Helping come up with names for things.
Puns, surprisingly.
Finding and fixing bugs in code.
Generating test data, and examples generally.
Telling stories.
Showing examples of formal writing that you haven't encountered before - I've used it to help me see what things like grant applications and sales manuals look like.
Games - so many fun games you can play with it.
Those are all examples of things I've used it for just in the past week.
I think something like automatically generating dialog for characters inside of video games in a way that's more dynamic and scalable than the clearly hardcoded dialog we have today.
People have mentioned using it to cheat on high school level English papers and blackhat SEO spam / fake social media accounts.
I had trouble parsing this sentence.
And why doesn't the abstract provide at least some basic explanation of the title?
The last line of the abstract does seem to be relevant:
> This work consists on (sic) an attempt to describe in a concise way the main models are (sic) sectors that are affected by generative AI and to provide a taxonomy of the main generative models published recently.
I think the authors may have meant something like:
"This work presents a concise overview of generative AI models and their application areas, and provides a taxonomy of recently published generative models."
They could have gone on to explain that the models are categorized by input and output type.
>This work concisely describes the main sectors affected by generative AI, along with a taxonomy of recently published models.
The AI/ML space moves so fast that this review is already outdated.
However, that is mostly do the effort required to change course. If we reach a point where it's easy enough to regenerate everything from scratch, will it be so important to correctly plan ahead?
Determinism still requires planning.
(this response generated by GPT ;-) )