OpenAI releases larger GPT-2 model
openai.com
openai.com
I'm working on tools to streamline GPT-2 text generation: I'm currently porting the code above to gpt-2-simple (https://github.com/minimaxir/gpt-2-simple) to allow easy finetuning/generation, and am also working on a way to quickly build an API/client for easily deploying GPT-2 to production and generating text at scale and cost effectively. Even the 117M model, managing CPU and RAM performance is tricky.
But given the incredible results of just the 117M model (e.g. Hacker News titles from a retrained 117M model: https://github.com/minimaxir/hacker-news-gpt-2), I'm eager to put the 345M model through its paces.
If it didn't set the flag, it gets messy: https://twitter.com/minimaxir/status/1124742105421631488
Wow, thats great. “The Bullshit Bubble” “Fuck you, Bootstrap” “We should give up on America” - they’re practically comedy, yet very believable too.
Though "Hacker makes an infinitely scaleable cup of coffee" gives it a good run for its money.
(I have an idea for a more proper approach)
[1] https://www.reddit.com/r/MachineLearning/comments/bmn0og/p_l...
Very reminiscent of headlines from The Day Today: https://youtu.be/wdEcO8_2Kl8
Hiring technical debt (or "unsortable overtime")
How do you hire technical debt??
Edit: another one-
“In 2009, Africa power creation was switched on for the Google Earth Darth Vader Imperial Warplane Propaganda”
I guess Google is diversifying :P
-Hiring -> Software Developer
-Software Developer -> Code
-Code -> Technical Debt
You only have to hire a Software Developer :)
There's so many nefarious things you could do with this...
Generate fake news. spam google. etc.
One valuable use could be to generate comedy and parody.
You could also make it to sabotage others too. You could set it lose on nazi forums and have them argue with bots constantly.
Just because lies and propaganda have always been around, it doesn’t mean they were never a problem.
Consider the issue of fake reviews: sometimes fake reviews are really obvious. Sometimes they aren't. Often the best way to pick out the fakes is to analyze all of the other reviews the user has written. That is going to become harder.
For very technical topics, where the reader comes with a strong background knowledge in that topic, picking out the fake material isn't too difficult. I suspect for the hazier things where the writers are more or less stating opinions, like politics, it is going to be incredibly difficult (for a human reader) to separate the bots from real people.
At a glance it's hard to see if the 345M model is "better" at the moment (it's all qualitative), but I'll be doing more testing. Unfortunately, the 345M model might be slightly too resource intensive for the API/client use case I wanted to do, so I'll likely be sticking with the 117M model for now.
Wow, this sounded so believable I had to check if this exists or not. It seems it doesn't exist, yet :).
The generated titles are great! You can put them into hncynic (https://github.com/leod/hncynic) to get closer to a fully generated HN experience.
If I had to compare them, I'd say that accumulation is about working on a minibatch datapoint by datapoint and faking being able to run an entire large minibatch in a single shot, while checkpointing is about working on a model layer by layer and faking being able to run an entire model in a single shot.
The problem with GPT-2-335M and why nshepperd had to mess with gradient checkpointing is that the GPT-2-335M model will literally not fit in your standard 11GB GPU (and from Twitter comments about people trying it on the new 16GB Google Colab instances, it's unclear if 16GB would be enough either!). You can't even run minibatch n=1. It doesn't fit. It OOMs.
The model itself is only a gigabyte or so, the problem is that the self-attention layers, when run, use up a huge amount of memory for their intermediate steps, which must be stored in order to trace everything backwards through each step for the backprop part of training.
(Right now I believe nshepperd's code punts on doing gradient accumulation simultaneous with gradient checkpointing, so we've just been reducing the learning rate, which is sort of similar to faking large minibatches with gradient accumulation.)
Fortunately, because the self-attention layers are so small and cheap to compute, they work well with gradient checkpointing. They're cheap to recompute on the fly, so it's more important to save memory and allow training at all. (This is also how OpenAI is training the Sparse Transformers which are enormous; they haven't said either way, but I assume this is how they trained the larger GPT-2s like the 1.5b parameter version, because I can't imagine what hardware would fit even a single GPT-2 1.5b without tricks.)
If I understand nshepperd's code correctly, it uses a constant and small learning rate. Do you know if this works better than the learning rate schedule that is usually used for Transformer models (https://www.tensorflow.org/alpha/tutorials/text/transformer_...)?
Just to pick your brain (I’m an ML noob and it triggered an idea in my head). I was wondering if GPT-2 could be used to generate small paraphrases for an input sentence for search suggestions? e.g “I’m going to the mall today” -> “I will be going to mall today”.
Plus, it's not as if this thing is more "potentially harmful" than, well... basically anything of use, like electricity, internet, fire, less perfect language models. In fact, it isn't even anything new, it's just (possibly) less broken language model than what we already have.
Admittedly, it would be quite problematic to use the full model with today's mainstream GPUs, so I'm not that much saddened by them hoarding on it. It just seems curious to me.
EDIT: Done! It takes about 3x longer than it did before to generate a response, so if you try it be very patient. Also, I made this in three hours so I wouldn't be surprised if it goes up in flames at some point.
Tried with this input: "what is your purpose as a bot?" Response: "There are a handful of examples available out there that seem fairly plausible as a good initial goal. These examples might get you started: A bot with a purpose to collect information (and I'm going to assume that you have to have an intent, so I'll be using that term and not the verb for my example). Or a bot with a purpose to help out a customer that wants something delivered to him/her by mistake. Or a bot that wants to promote your site to the entire universe of bots. For example, here in New Zealand, that's a fairly common use of botnets as far as I can see: How do you know I have enough information to set it up? Let's say I have an account with the website at /r/frugal, what if you are already running bots that want what I need? How do you know these bots would actually be helpful? One answer I see is"
Pretty cool stuff, but I see your description as "semi-coherent gibberish" fits really well. I could see something like this get pretty interesting in the future, where you, for instance, could maybe populate a discord server or twitter chat with semi plausible noise to fake popularity.
hugging face has a simple implementation that augments bert in this manner and you can see the code there. their bertqa model get like an 84 F1 on squad 1.1 which really strong performance. you can augment the thier gpt2 implementation similarly.
https://openai.com/blog/better-language-models/
It's fascinating to think an AI wrote that story...
Has anyone said when they are intending to release the full model? There's likely a lot of positive applications of this technology as well.
At the same time, the risk is lower that a different team develops similar capabilities and widely deploy it without anyone even knowing that what they do is possible.
Wouldn't be surprised if this pattern is followed even more in the future. It's an acknowledgement that there are many angry cats that will (almost unstoppably) be let out of the bag in the next decades, but attempting at releasing the cats in the most responsible manner possible.
There has. BERT has come out and is better.
With the main concerns being troll army/fake news type stuff, I don't think this makes a difference. We seem pretty sure there are state level actors behind a lot of that stuff, and I think it would be silly to believe they can't recreate something at the level of GPT-2, especially with the underlying principles out there and understood, competitors like BERT available, etc.
I think their heart is in the right place, but also incredibly naive.
The net result of these advanced forms of signal processing will be negative. Nobody has come forward to prove that they will benefit society on the whole or even that they are safe. But anyone who raises concern is shouted down and called names like “alarmist” and “Luddite.”
These companies are playing with fire, and the whole world stands to be burned. Wake the fuck up.
This isn’t something that can be built and tested in isolation like other things we are familiar with. Training these models is not an exact science. Nothing about ai is an exact science. Progress only comes with trial and error. And each trial requires huge compute resources; at least for the most capable and dangerous models. It can’t be done in your basement. Not without significant effort and drawing attention to yourself. Could we sense whenever someone was trying to do it? Could we form a global coalition to stop every attempt? That brings us to the next thing.
What you are doing is the following: we are both in a car that is about to roll off a cliff. I propose that we try pressing the brakes. You respond by saying that, geez it looks like we probably wouldn’t stop in time — we are going awfully fast and it probably wouldn’t work to press the brakes so why even try? Let’s just brace our heads and hope the impact doesn’t kill us.
Obviously the better thing to do is to try and press the brakes. Even if you aren’t sure if you can stop in time.
Imagine a thing "we should not do", let it be creation of really powerful language model, genetic engineering in order to produce smarter, stronger children, or whatever. Whatever you are opposing to, really. The simple fact is that if you (and by "you" I mean any entity you associate yourself with, be it literally you, or your company, or a group of researchers in your country, or the government of your country) can do something, anybody can. Even if it will be a couple of years later. You are not unique, nor alone. You can stop "yourselves" as long as you want, but there are other people, companies, research groups, governments, and they don't give a fuck about what you think "we" should do.
So, in the end, the car really doesn't have any brakes. Even if it truly means the end of it all, it's just unfortunate, but really, really unavoidable.
OpenAI suffered a ton of blowback for not just releasing the full model from the start. You can read their intitial blog post [1] looking particularly at the sections for Policy Implications and Release Strategy. I would also highly recommend you listen to Lex Fridman's podcast with Greg Brockman [2] to hear their rationale about the recent org changes at Open AI.
Obviously you can posit that everything they say is bullshit and they are only after almighty dollars. I can't prove it's not true and personally believe there is at least a kernel of truth to it, but we live in a messy world and finding imperfect allies is generally better than having none at all.
Also, yeah, you could do it in your basement without being detected.
You can do it in the cloud, too, a lot of us here have the skills and resources to do it, but why spend time on this toy instead of another one, especially if it's going to cost money we could spend on something more fun?
No. What I am saying is that it's impossible to centrally control the actions of 7.7 billion free humans. A lot of them will disagree with your position (and any other position as well).
By trying to "put a stop to it" in a central manner, you are only making it harder (but not impossible) for some subset to learn about this phenomenon, to improve on it and to understand its strengths and weaknesses.
I am unconvinced by your argument that we could reliably detect, much less stop, attempts to train a large and useful machine learning model.
OpenAI is a non-profit. They are not looking to get acquired.
Saying hyperbolic things like "the whole world stands to be burned" is silly. They aren't giving everyone a nuke.
Things like deep fakes and synthesized speech are much much worse, since you can make any politician say anything you want, and an average person wouldn't be able to tell it's fake.
>These companies are playing with fire, and the whole world stands to be burned. Wake the fuck up.
Your tone and arrogance is very unwelcome.