GPT-2: 6-Month Follow-Up
openai.com
openai.com
"The Middlebury Institute of International Studies Center on Terrorism, Extremism, and Counterterrorism (CTEC) is exploring how GPT-2 could be misused by terrorists and extremists online."
"The University of Oregon is developing a series of “bias probes” to analyze bias within GPT-2."
But apparently no university studies the social and economic impact of using terabytes of public data to train algorithms that for all practical reasons end up being inaccessible to an average person.
If things go on the way they're going right now, in 20 years millions of people will be "mechanical turked". Most of information processing tools will be mediated exclusively through companies like Google and Amazon. They will be less like normal tools (e.g. word processors) and more like systems you have to be a part of. Can you imagine the levels of inequality involved? The hyper-centralization of power? This is the foremost challenge presented by AI, not some hypothetical nonsense involving terrorists using a text generator.
And it's not like there aren't any solutions. Douglas Engelbart, for example, pointed out a great way of introducing technology into society without screwing most of the society over:
http://dougengelbart.org/content/view/138
We kind of followed his vision for a while, with good results, but AI seems to be going in an entirely different direction.
Seems accessible to me
For a lot of nonstate orgs with sophisticated propaganda arms, an ideologically cohesive text generation capability would be a huge advantage in scaling up info ops. We are looking to measure whether or not GPT-2 or other neural text generators are useful for this, or if that risk is, as you say, nonsense.
I think it's pretty clear that terrorists or any other bad actor will find great value & utility in this tech. The article from OpenAI says 'Humans can be convinced by synthetic text.' & research at Cornell found people find it almost as convincing as New York Times articles. I would be interested in learning about the methods you guys are using to determine of it? I wonder how that could be measured?
So let's assume the answer is 'YES! This technology is dangerous". The Middlebury program, Cornell, and more and more universities and research groups find the same thing. Then what will the recommendations be? Certainly not to release it into the wild. I think they will be to keep it locked up. To keep it in the hands of a few large and powerful companies, with the resources to 'manage' such a thing.
This seems to be what the original comment is trying to illustrate, and I think it's an interesting point to consider the implications of long term. The tech exists now. There is no going back. So is it worse to let it out of the box, or to let a but a few have control over it?
I suspect they will release the full model in time. It's already trending in that direction.
Like I replied to the other poster, I strongly support open-sourcing tech. Centralization of tech like this helps exacerbate the problem: state and sophisticated nonstate groups have the resources to develop it indigenously, while the public can't dig into it and start developing a set of norms and best practices to approach detection and mitigation.
"Anyone" includes researchers. That's part of the deal. Yes, they benefit, but you aren't harmed. That's zero-sum thinking.
Perhaps what the poster above you is saying is that there is a continuum of information, some more personal and sensitive, like your current location, and some less personal, like the Wikipedia article on Elephants.
Taking data about specific humans, (or humans in general), and turning it into code that has predictive power seems like a different type of power than the knowledge given by an encyclopedia.
One could argue that I actually started work on misinformation though Engelbart. I was working on various projects specifically to achieve a vision like his, and it is still my guiding light. But it turns out that some technologies (initially focusing on Facebook and YouTube's engagement dynamics...), without appropriate checks and balances, are sort of anti-augmentation of intellect. They give an asymmetric advantage to those who are trying to weaken our intellects. So I ended up dropping those projects to attempt to address urgent misinformation issues in May of 2016.
Going back to GPT-2 and research release, I and my co-author recently went deep into understanding the types of risks and tradeoffs in a recent paper. You can see the summary here: https://medium.com/@aviv/reducing-malicious-use-of-synthetic... or go directly to arXiv https://arxiv.org/pdf/1907.11274.pdf . The goal of our paper is specifically to go past the angry invective of the "here is the most important problem" (in your case, "AI inequality") and actually dive into the weeds of threat models and tradeoffs.
Also, with so many people concerned about disinformation, where is research on tools that would empower individual users to process information in better ways and make sounded judgements?
The problem with this approach is that the users most likely to be malicious users of GPT2 are state actors. China, for example, already spends millions on an immense propaganda factory. Money is not a serious obstacle for a state. Given that other research entities are, by the sound of things, already far along with development of similar models it seems unlikely that China and the US don't already have functional models internally.
On the other hand, legitimate business and research is clearly hamstrung by withholding the full model. What we have is the maximum degree of inconvenience and the minimum degree of security. It feels almost perfectly analogous to ban on liquids in airports. The motivation for that ban was that existing security measures couldn't detect liquids, but simply announcing a ban was to be enforced didn't change the fact that liquids were undetectable. Instead millions of travelers were pointlessly inconvenienced at great cost.
Release the kraken already!
This decision, as well as our discussion of it, is an experiment: while we are not sure that it is the right decision today, we believe that the AI community will eventually need to tackle the issue of publication norms in a thoughtful way in certain research areas.
This release approach is an experiment used to force the conversation around a release strategy before we actually and unambiguously need it.
Of course, the default output from the model might be sufficient, although it’ll take twice as long to generate text compared to the 345M which is slow even on a GPU.
How exactly the large GPT-2 models are deployed is a mystery I really wish was open-sourced more.
>What would be the largest model one could train across 2x 2080Ti?
>~800M gpt2. this is largely due to the memory required to House parameters + optimizer states. If one uses a smaller optimizer than Adam training something larger should be possible. Make sure to turn on activation checkpointing with —checkpoint-activations
>We do provide training code that should work out of the box for gpt2 117M/345M
TalkToTransformer.com uses preemptible P4 GPUs on Google Kubernetes Engine. Changing the number of workers and automatically restarting them when they're preempted is easy with Kubernetes.
To provide outputs incrementally rather than waiting for the entire sequence to be generated, I open a websocket to a a worker and have it do a few tokens at a time, sending the output back as it goes. GPT-2 tokens can end partway through a multi-byte character, so to make this work you need to send the raw UTF-8 bytes to the browser and then have it concatenate them _before_ decoding the string.
While my workers can batch requests from multiple users, the modest increase in performance is probably not worth the complexity in most cases.
From my observations most SOTAs come from chineses researchers by far, followed by deepmind.
BTW isn't that a sad truth that not even one of all major AI actors has a draft of an AGI architecture, something comparable to CYC or opencog. https://wiki.opencog.org/w/CogPrime_Overview
Two other observations I would like to share: Many important NLP tasks have almost nobody publicly working on them it seems, on paperswithcode.com or NLP-progress (from github) some tasks have only one or two papers... And many others have not evolved since 2016. Most of the time it seems trivial to beat the old state of the art, just use BERT or XLnet on a task where nobody applied it before and hop, free state of the art for you! Yet researchers don't seems to chase those low hanging, high returns fruits. Also researchers seems to work a lot in isolation, many new generic improvements like new optimizers (RAdam for example) and new activation functions (Swish) allow to beat most of older state of the art on almost all task just by using them. Yet researchers will take years before using them because of an absurd inertia. Also unlike an open source program, BERT and XLnet have very low response and activity on github despite major open issues... </rant>
Well, then perhaps you should go work on them, instead of ranting here.
There's nothing to fix. People work on what they want to work on. Things that seem important to you are not important to me, and the opposite. I'm OK with that.
I'm not sure. I rarely have to do that explicitly in my head. Perhaps a model should learn to infer/guess them implicitly, from context, just like I do.
what exactly is more important?
In my opinion, having a world model (for common sense) and situational awareness (e.g. through sensor fusion, or from prior conversational history, or using some externally supplied conditioning) would be far more important.
GPT-2 does not generate meaningful sentences or only by accident
You think adding POS tags would help it generate meaningful sentences?
In my opinion, having a world model (for common sense) and situational awareness (e.g. through sensor fusion, or from prior conversational history, or using some externally supplied conditioning) would be far more important. Haha you basically want a general intelligence (AGI), I want it too! And not enough persons works on "architecting" such a thing. Opencog may interest you a lot then. But the reality is many other "simpler" tasks are needed to make this happen.
having a world model (for common sense) is an NLP task There are some interesting results https://github.com/sebastianruder/NLP-progress/blob/master/e... OpenAI does not work on this task sadly, at least for now.
You think adding POS tags would help it generate meaningful sentences? I would be clearly insufficient yet necessary. I believe they already use internally a POS tagger and a dependency parser.
Interesting. Where did you see that?
GPT-2 has not (yet) been trained to predict POS tags to my knowledge, nor BERT, or ernie 2 or xlnet has, but I think they have great potential to improve POS accuracy.
It's useful to view the GitHub projects for these models as reference implementations. They're intended to provide a roadmap for reproducing the research and to aid in implementing production libraries.
Regarding the latter, take a look at the work by HuggingFace, the Flair project, Spark-NLP and others.
"While GPT-2 has gained mindshare for its generative capabilities, BERT and other pre-trained Transformer Encoder models are used for production workloads" You rephrased my point pretty well, while openAI search for "fun" tasks, deepmind and others allow progress on real world tasks.
You use BERT which is nice but do you consider using it's successor: XLnet?
"take a look at the work by HuggingFace, the Flair project, Spark-NLP and others." I was aware of Flair (from Zalando) but thank you for Huggingface and Spark-NLP, I will take a look!
Writing Bloomberg's "market wrap" articles.
And earlier this month they released a local version: https://www.tabnine.com/blog/local/.
If it was a high-return fruit somebody would be doing it. Not necessarily publishing papers about it or trying to beat useless artificial benchmarks on it.
Not necessarily publishing papers about it Yes.
trying to beat useless artificial benchmarks on it. Wtf is this bullshit? AI benchmarcks are what direct progress in AI and allow to quantify it. And they are less and less artificial and more and more real world: E.g quora, reddit, Wikipedia and Facebook datasets.
> AI benchmarcks are what direct progress in AI
Sadly this is largely true.
The AI benchmarks + culture around it are the bullshit.
What actually moves forward the field of AI is:
- accessible
- reproducible
- comprehensible
results done with some thought and reasoning which is explained well, published well, and justified by more than some #$!& "our F1 score went up by 2 therefore our approach makes sense" bullshit.
AI benchmarks have done as much to retard progress in AI as they have to promote it.
Current AI benchmark top scores are gamification for big companies to waste even more resources running algorithms they can't explain. They are not machine learning, they are machine pissing contests.
https://medium.com/@VictorBanev/interrogating-gpt-2-345m-aaf...
Not sure why open AI doesn't do this themselves. That fully aligns with their stated mission.
I found that it might be very effective. I have the test at
https://docs.google.com/forms/d/1p7tlobl5y5plBCu_enK4KawR7B8...
I got the information from trumptwitterarchive.com
I also explored creating a system that could recognize fake tweets from real ones and I believe I got 94% accuracy. It was a Bayes classifier but I think I have to double check my work.
This seems naive but I think it's a misdirection. Of course the model will have malicious users. Propaganda teams started testing its integration as soon as it was released. It's likely that OpenAI is counting on this for insights into HOW the model can be used maliciously. It's also possible that the model results have inherent trackable markers and OpenAI can later say that X% of social media posts were made using this model.
So what are the positive applications, aside from prettifying data like sports and weather reports?
Even with Skyrim's 800+ books, you frequently ran into the same book. Imagine libraries filled with plausible text that hides nuggets of lore seeded by developers. Along with more realistic text-to-speech this can allow games to support a large diversity of NPCs that have true radiant dialogue and sound more realistic than "I saw a mudcrab the other day".
With some modifications, I think models like this can outweigh even their nefarious applications:
Defense against text decomposition analysis. The model can be used to obfuscate writing patterns that can reveal a person's identity, either by randomizing form or standardizing it. Take your post and run it through the formatter to get the same idea and intent, but in a style that can't be traced to your other writing. Or you reform it into style of Ernest Hemmingway, like thousands of others.
Realtime plausible deniability encryption. Messages in a monitored chat can look like mundane conversation but contain encrypted messages. This would require the model accept seeds and work partially in reverse to diff two sets of text to reveal the hidden message.
In it's current form it doesn't look like it can do any of those things, but there's the potential.
As an example, look here by Nvidia https://devblogs.nvidia.com/training-bert-with-gpus/ who made GPT-2 8B, which is ~5 times as large as GPT-2.
But some of you statements seems incorrect: Ernie 2.0 being stronger than xlnet XLnet is the neural net with the biggest number of first places on benchmarck leaderboards. Cf: https://paperswithcode.com/paper/xlnet-generalized-autoregre... While ernie 2.0 has currently 0 first place on paperswithcode.com https://paperswithcode.com/paper/ernie-20-a-continual-pre-tr...
xlnet trained on 10x more data than original BERT No, I've read on a github issue of xlnet that xlnet base is same size as bert base and xlnet large is same size as bert large. (I don't know for ernie 2)
Well your point on finetuned bert vs non finetuned xlnet is interesting. ROBERTA is so fine tuned it beat XLnet on some tasks. But generally xlnet non finetuned beat BERT finetuned and there are more and more xlnet finetuned each week. But your point does apply for Roberta, and for the few tasks where bert as been applyed but xlnet hasn't yet.
It's not about the size of the model, but the training data. If you read the XLNet paper https://arxiv.org/pdf/1906.08237.pdf they clearly state in section 3.1:
"Following BERT [10], we use the BooksCorpus [41] and English Wikipedia as part of our pretraining data, which have 13GB plain text combined. In addition, we include Giga5 (16GB text) [23], ClueWeb 2012-B (extended from [5]), and Common Crawl [6] for pretraining. We use heuristics to aggressively filter out short or low-quality articles for ClueWeb 2012-B and Common Crawl, which results in 19GB and 78GB text respectively. After tokenization with SentencePiece [16], we obtain 2.78B, 1.09B, 4.75B, 4.30B, and 19.97B subword pieces for Wikipedia, BooksCorpus, Giga5, ClueWeb, and Common Crawl respectively, which are 32.89B in total"
If you compare to BERT paper https://arxiv.org/pdf/1810.04805.pdf training data for "pretraining data" section in also section 3.1:
"Pre-training data The pre-training procedure largely follows the existing literature on language model pre-training. For the pre-training corpus we use the BooksCorpus (800M words) (Zhu et al., 2015) and English Wikipedia (2,500M words). For Wikipedia we extract only the text passages and ignore lists, tables, and headers. It is critical to use a document-level corpus rather than a shuffled sentence-level corpus such as the Billion Word Benchmark (Chelba et al., 2013) in order to extract long contiguous sequences."
So 32.89B words for XLNet vs 3.3B words for BERT.
We've also run fine tuning experiments supplementing additional private medical corpus (~10B words) and felt starting from clinical-bert was better than xlnet (for our rather specific use cases).
EDIT: Looks like https://talktotransformer.com/ already uses the 774M one!