Turing-NLG: A 17B-parameter language model
microsoft.com
microsoft.com
More than that, I think NLP will unlock new ways of interacting with computers. Computers will be able to handle the ambiguity of human language, transcending their rigid “only do exactly what you tell them” models of the world.
Edit:
Adding this to give more technical context. I think most people don’t know where the line is currently between what possible, and what’s not, but also what we are on the cusp of. And we are on the cusp of a lot.
A quick explanation of one area is here:
Basically, transformer models are the best for NLP. They use something called attention based mechanisms, which allows the model to draw correlations between pieces of text/tokens that are far apart. The issue is that this is an O(n^2) operation. So the model is bounded by the context window, which is currently mostly at 512 tokens, and is thus, bounded in how much it can understand. Recent innovations, and further study, will broaden the context window, and thus unlock better reading comprehension and context understanding. For instance, the ability to answer a question using a piece of text is mostly stuck at just finding one paragraph. The future will see models that can find multiple different paragraphs, understand how they relate, pull the relevant information, and synthesize it. This sounds like a minor step forwards, but its important. This will unlock better conversational abilities, but also, better ways to understand how different pieces of textual information relate. The scattershot of information across the internet can go away. Computers can better understand context to act on human intention through language, unlocking the ability to handle ambiguity. This will change the internet.
Again to empathize, these models only started showing up in 2017! The progress has been rapid.
For the most part they aren't successful because the "AI" isn't smart enough to have a goal in mind, so they end up just monkey-cheesing everything. Pasting together snippets in ways that are usually grammatically correct but make no sense.
Part of this is "underlying meaning" is an intuitive way to describe things but whatever is underlying here is more tenuous than a classical logic/GOFAI model of the world but more "solid" than a long, clever stream of associations.
Let's say we have a goal, "evaluate persons impression of a particular book/topic/etc".
So the goal would be to have a conversation on this and related topics that would (re)construct person's impression.
Hence my question if there are any publications/articles that explored that?
The cold hard truth about statistical (and by extension, deep) NLP is that it's just a fancy way of counting numbers mostly. The only way to get to _real_ language understanding is AGI, and _nobody_ is working on that. You fundamentally cannot interact comfortably with a human if your system does not have probabilistic, contextualized cognition, and can't incorporate knowledge about the world.
I’m not saying these NLP methods will be some kind of AI, just that they will produce products, content, and ways of interacting with the world that are categorically different from what we have seen in the past.
For instance, question and answering tasks have only recently been able to:
Find an answer in a text document that spans multiple non contiguous paragraphs
Understand context across a whole book.
The context window of current nlp is stuck at 512 tokens, mostly because of computational complexity. This has been broken just recently by the reformer model. Which is a primitive, early way to get around the computation costs of attention mechanisms.
Just wait. The ideas are there. They just take time to refine.
The fundamental limitation of these "optimization tools" as you call them - they don't have any common sense, and any way to query an external source of information (e.g. wikipedia), or ask a human to clarify.
Another big problem is we don't have any way to do quality filtering on the outputs. From my experiments with GPT-2, it produces one interesting paragraph of text out of 20 - if you squint at it really hard. And most of those 20 don't make much sense at all.
So no, the existing ideas are definitely not enough. Maybe some novel hybrid of symbolic AI with statistical optimization will lead to a breakthrough. This one does not strike me as anything other than "let's use moar weights!!"
What do you mean by this? Of course they do, learned from their training data. For example, here is quote from conversation 38 of https://github.com/google-research/google-research/blob/mast...
Human: Do you like Korean food in general? Meena: It's okay. I like beef bulgogi, but I'm not a huge fan of kimchi.
It seems to me Meena "knows" bulgogi and kimchi are Korean foods. Isn't that common sense? If it isn't, what do you mean by "common sense"?
Human: Do you like Korean food in general?
Meena: It's okay. I like beef bulgogi, but I'm not a huge fan of kimchi.
Human: Ok what should I shop for ?
Meena : You've got almost everything but you need a pear, the steak and some ginger.
The problem with language models as commonsense is that they are collections of patterns and associations, and that they don't have inference models or solvers - unlike my dog for example!
A more relevant analogy might be a talking parrot :)
Lin, B. Y., Chen, X., Chen, J., & Ren, X. (2019). KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning. 2822–2832. https://doi.org/10.18653/v1/d19-1282
Liu, W., Zhou, P., Zhao, Z., Wang, Z., Ju, Q., Deng, H., & Wang, P. (2019). K-BERT: Enabling Language Representation with Knowledge Graph. Retrieved from http://arxiv.org/abs/1909.07606
Trinh, T. H., & Le, Q. V. (2019). Do Language Models Have Common Sense? Iclr, 1–12.
Ostendorff, M., Bourgonje, P., Berger, M., Moreno-Schneider, J., Rehm, G., & Gipp, B. (2019). Enriching BERT with Knowledge Graph Embeddings for Document Classification. Retrieved from http://arxiv.org/abs/1909.08402
Let's take a question answering example. Take just about any recent deep learning paper and try to answer detailed, higher level questions against it. To use a concrete example, take MobileNet V3 paper and ask your system "do I put activation before or after squeeze and excitation" (correct answer is "before"), or "do I need a bias in squeeze and excitation module" (correct answer is "it depends on the task"). You won't be able to, because a lot of things are just _assumed_, just like in any other realistic example of text written for human consumption. The facts are encoded externally as information about the world, and they're so fine grained and contextual, that we don't even know how to begin incorporating them into the answers, let alone do so contextually and probabilistically, like human mind does.
I’m optimistic because I believe that the contextual information that you are describing, is already there in the vast expanse of the internet.
But I will also add, I think none of this will spawn AI, just that it will spawn new technologies that are categorically different.
That is a rather contentious claim.
I don't think anyone familiar with the area thought that ConvNets will give us AGI.
However, their effect has been huge! It's hard to overstate this. Computer vision used to be a small niche topic, with tons of effort required to get something working even on simple images. The quality of today's ConvNet predictions is way beyond anybody's imagination in around 2010. Models built around that time were like a house of cards. Extremely carefully crafted for specific scenarios, where moving one threshold a bit would destroy your output.
Also sometimes in medical imaging the conditions are very different from actual practice. For example the doctor may be worse than the convnet on certain types of low-quality, low-dynamic range images that someone preprocessed in a particular way. But sure, in the medical field some error prone, boring counting tasks and spot-the-cancer-in-your 200th-image-today, the machine can perform actually better.
But what tasks specifically do you have in mind?
>Understand context across a whole book.
to have truly much utility when it comes to context a model doesn't just need to correlate information across some text, it also needs general knowledge and understanding so that it can produce knowledge which is only implicit or not even present in the text itself.
You can make the transformers as large as you want, NLP models still fundamentally suck at answer trivial questions like "If Pete was alive in 2000 and alive in 2050, was he alive in 2025?"
Basically, transformer models are the best for NLP. They use something called attention based mechanisms, which allows the model to draw correlations between pieces of text/tokens that are far apart. The issue is that this is an O(n^2) operation. So the model is bounded by the context window, which is currently mostly at 512 tokens, and is thus, bounded in how much it can understand.
Recent innovations, and further study, will broaden the context window, and thus unlock better reading comprehension and context understanding.
For instance, the ability to answer a question using a piece of text is mostly stuck at just finding one paragraph. The future will see models that can find multiple different paragraphs, understand how they relate, pull the relevant information, and synthesize it. This sounds like a minor step forwards, but its important.
This will unlock better conversational abilities, but also, better ways to understand how different pieces of textual information relate. The scattershot of information across the internet can go away. Computers can better understand context to act on human intention through language, unlocking the ability to handle ambiguity. This will change the internet.
I'm happy to be wrong about this, but I'm not seeing any discussion about the safety and security of using these systems. And if it's not even being discussed, we can be sure nothing's actually being done about it. Selling promises of active-agent computers interpreting human intent and summarizing information from the Internet without addressing this concern is irresponsible at this point.
And all of _that_ is just to interpret sincere, honest attempts at communication. Can it safely and appropriately handle humor, irony, and sarcasm? What about coordinated malicious attacks?
This is similar to how computer vision models work great in most cases, but to build a self-driving car you still need all the other components that do path planning, predicting what the cars around you will do based on the state of the environment, etc.
I agree, there needs to be a way to represent relationships between information. I personally don't think knowledge graphs will be the ones to do it, not because they dont work, but because of how imperfect they are, in the data quality sense.
See this paper here:
"DIFFERENTIABLE REASONING OVER A VIRTUAL KNOWLEDGE BASE" https://openreview.net/pdf?id=SJxstlHFPH
Which is a recent effort, among many, by google research to build a model that can view a document as a knowledge graph, instead of explicitly tying pieces of the document to the graph, the idea to is to create a graph from the document. This is paper is a bit different from that, they do input a knowledge graph for training, but I think the idea and track of where they are headed has a ton of room to evolve. The trick is that transformer models have unlocked the ability to understand the text, so all of this "quasi knowledge graph extraction" that i was just explaining, is only recently possible! There's no research on it, because the baseline understanding of tokens has been too primitive. This is why there is so much room to grow, BERT has unlocked new methods, it can be used as a base for a ton of new NLP.
Just to emphasize again, I'm not saying what I outlined above will be a good way to do it, just that ideas like this could only be tested recently. There's a million new ways to spin this problem.
Such as google's duplex?
GPT uses 1024 tokens context window which does work out to a fair amount (given the massive vocab of 50k+ which means a token can be more than a word), though of course it's pretty limited.
Google's recent Reformer[0] allows you to do attention much more cheaply and I'm currently training a Reformer model that isn't quite as big but has a context of ~64k tokens (though a much smaller vocab). I'm not completely sure if this is the solution but it looks like a step in that direction and so far the model is doing pretty good (I also plan to post the weights when I'm finished, though I am not sure if Google don't just plan to do that themselves).
I am somewhat disappointed they went with just 1024 for this model, too though.
0. https://ai.googleblog.com/2020/01/reformer-efficient-transfo...
I expect there to be an improvement like there was for
BERT -> Albert
So, are reasonable examples now of these models allowing semantic context? So, far, what I have seen is generated text where the lack of understanding takes three paragraphs to become obvious rather than one.
Human language is this marvelous framework involving symbols associating with other symbols as well as to well-known and vaguely-guessed facts about the world.
Human relations are very robust and, for example, two people can have a longish conversations where at the end, they realize they're talking about two different people (or different days or events). But in those circumstances, they can correct and adjust. "Solid" understanding is there but it's under a lot of layers of social cues and protocols and multiple meanings.
This is about where I am stuck. I'll start believing that we truly are on the cusp of a revolution as soon as I see Google Translate reliably knowing when to translate "home" into French as "domicile", "foyer", something those lines, or as "accueil."
Right now it seems to very frequently choose "accueil", which is generally wrong, except when you're talking about websites and software user interfaces. That it's biased so strongly toward that error speaks volumes about how critical semantics are to sorting out natural language, and also about how bad current NLP systems are at dealing with semantics.
Isn't that basically the same as the Winograd problem?
People often assume a very benign, civilized environment for AI. In ordinary life, human beings (on the internet or off) take an adversarial approach to other people modeling the logical framework behind an utterance. They either subvert it for amusement (trolling or comedy) or profit (politics, propaganda, sales), and a tremendous amount of effort goes into it, much of which is very effective.
As people have observed, it can be very easy to convince a human that a machine is intelligent, as with ELIZA. But if you violate the presumptions of trust that a machine is designed with, it's going to be surprisingly vulnerable. To be superior to humans, a machine would have to be able to fend off an intelligent person trying to undermine it, not just work when it is spoon-fed.
I'm going home -> Je rentre a la maison
Home sweet home -> La douceur du foyer (the translation is weird but the word foyer was expected)
I feel at home -> Je me sens chez moi (this one is particularly good, it didn't translate the word home directly)
Can you share your exemples where it fails?
As exciting as this sounds, I can't help but feel that given -we- haven't figured out how to handle the ambiguity of human language, I'm not convinced a computer attempting to is really markedly better for many use cases than requiring exactness. But operating at a human level of 'understanding', and being broadly accessible, may be enough to change the world. Hopefully for the better.
Maybe, maybe not. IMO, we don’t know what problem we have to solve to get what you describe. We also don’t have a metric as to how far we are along the path towards that goal (do we need 100B parameters? A trillion?), nor do we have any idea as to whether the current approach can get us there.
Yes, this year. While transformers certainly present a breakthrough in the NLP community and certainly stir up the state-of-the-art again, I don't really see how you go from that to the "computers will understand us" conclusion to be honest. People said that during the word2vec stir up and what we got out of that was incremental results (which is not bad, it's in the nature of things really).
We can already build models that do everything you describe as single tasks, while that's exciting, it's not going to lead to the singularity. We've got a long way to go in terms of understanding models, making them computationally tractable, and making them do what we want in the first place without resorting to hoping that our unsupervised model learns something useful. It's likely that the attention mechanisms we see today will be a large part of that but I'm honestly a bit baffled at the "People are vastly underestimating the changes that are about to come from NLP." part. They're not, people already think that today's AI is magic, I don't think it is benefitial to reinforce that. Speech is nuanced, we're making good progress in many areas but we're not on the cusp of any revolutionary change in computational understanding really.
Remember Google's demo of AI reserving a spot at a barber shop? Yeah... that never happened, even though it was supposed to be any day now.
This is begging the question as to whether the model "understands" anything at all. And once you adopt a definition of "understanding" that isn't equivalent to "got a high score on some pointless academic challenge" the answer is a resounding "no." The whole enterprise of AGI hype is based on this equivocation of words like "understanding" and "intelligence." We use a very restricted definition in proving that the tech is smart, and then switch out our restricted definition for the colloquial one when the audience isn't looking.
> This will unlock better conversational abilities
Shouldn't be hard given that as it stands there are none, except for creating a human-sounding slurry that is devoid of real content.
> but also, better ways to understand how different pieces of textual information relate
Is this a real need? What problem does this solve that forums + wikipedia + arxiv + google + a literate human hasn't already?
If you have time to discuss, DM me on twitter - I am @ralphbrooks.
You have several thousand documents in HTML, with roughly similar content (to a human) but not entirely consistent in the ordering of the sections, or the formatting, or the names, or the language and structure. But each document describes an entity of the same class (to a human) and very nearly all of them have a section that summarizes the document. They have a lot of parts in common, but they are fundamentally not designed to line up with a structure for data processing.
Is there any practical way to find that summary? Sure, something obvious that takes no time to script is to look for something like "Overview" but you quickly get bogged down in exceptions.
Maybe this seems very mundane and simpleminded, but it is the sort of thing that many people would assume is best done by a human. Is there anything current or in the near future that would be significantly easier than a human reading all of them?
* GPT & language generation models: Given some context (say a sentence), they can generate text to complete it, or to summarize it, etc. The task here is to actually write something.
Out of the box, given a sequence of n tokens, BERT returns a tensor of dimension (n_tokens, hidden_size) [1]. Where hidden size has no relationship with the vocabulary. You can then fine-tune a model on this representation to do various tasks, e.g. sentiment classification. Thus BERT is said to be a language representation model.
Out of the box, given a sequence, GPT-2 returns a distribution over the vocabulary [2] from which you can draw to find the most likely next word. Thus GPT-2 is said to be a language generation model.
You could of course play with the masking token of BERT call it recursively to force BERT to generate something, and you could chop off some layers of GPT-2 to get some representation of your input sequence, but I think that is a little past the original question.
[1] https://github.com/google-research/bert/blob/master/modeling...
[2] https://github.com/openai/gpt-2/blob/master/src/model.py#L17...
"BERT returns" is ambiguous here. During pretraining last layer is loggits for one hot vocab vector, the same as in GPT: https://github.com/google-research/bert/blob/master/run_pret...
Here’s a demo of BERT https://www.pragnakalp.com/demos/BERT-NLP-QnA-Demo/
Specifically for Transformers - any plans to train a big model with a bigger context window?
Not that this one isn't very impressive, of course.
Four days!
2) How long in years?
Three years!
The only applications I can think of for text generation are malevolent ones: I'm sure it would be great at generating spam sites which can fool Google's PageRank algorithms, and it seems like you could easily use it in an information warfare / astroturf setting where you could generate the illusion of consensus by arming a lot of bots with short, somewhat convincing opinions about a certain topic.
Is there something obvious I'm missing? It seems too imprecise to actually deliver meaningful information to an end-user, so I'm frankly baffled as to what its purpose is.
What’s the deal with these private demos? (GPT-2 was also essentially private). More importantly, why even announce the existence of a private demo to people who were not invited?
My take from the past few years is that we're 99% done with the visual cortex - convolutional nets can be trained to perform any visual task a human can in <100ms. Now I'm mostly convinced that GPT2 has solved the language cortex, and can babble as well as we will ever need it to. We just need a prefrontal cortex (symbolic processing / RL / whatever your pet theory is) to drive the components, which is a problem we have not even started to solve. I am 90% sure it is a different class of problem and we won't knock it out of the park in 5 years like the visual/language cortexes, but we can hope.
edit: it's possible cognition follows from language, which would be convenient. is GPT2 smarter than a dog? I don't think so but I could be wrong ¯\_(ツ)_/¯
While Markov chains sound like a schizophrenic, this sounds like a spacial disjointed notebook, as if somebody was trying to write in two places at the same time. "Today is Monday, it's Saturday night, I forgot to write to my dad and he only leaves the house for a couple of hours."
Yet, sometimes Markov chains turn out to be the most beautiful art form. His wife Jennifer Neil, whom he met at a barbecue and has been married to since 1998, attributes this creative process to the constant ups and downs in his old job.
"His numbers are just insane," says Jennifer Neil. "I don't know where he keeps them, but they're very mind boggling."
Somewhat disconnected from the actual
But the article is fascinating nevertheless. Not sure is alphago breakthrough.
Does anything like this exist?
Otherwise, you'd reach a word like 'and' and couldn't possibly follow it with a logical statement that follows on from the previous part.
My point being that these generation models should be conditioned on something more than just word history, like something they want/are instructed to express.
"I'm really starting to get worried about my Higgs Boson (HBN) after watching some videos on YouTube" [0]
"This repost and all of your posts are garbage." [1]
"It's the most random and unoriginal shit I've ever read." [2]
[0] https://www.reddit.com/r/SubSimulatorGPT2/comments/f1sqyh/an...
[1] https://www.reddit.com/r/SubSimulatorGPT2/comments/f1vfnv/wh...
[2] https://www.reddit.com/r/SubSimulatorGPT2/comments/f1o83a/fo...
GPT2 is still very random and quite stupid.
You start it with your love for your girlfriend as a context, she becomes a cam girl into hard core anal two paragraphs later. You start with religion, "Muslims must be exterminated". You start with software and you get a description of non existent hardware with instructions about how to setup a VPN in the middle. You start with news, and you can read than China supports the Islamic state.
That's cool because it has more context than Markov chains which usually have only 3 words of context, but it's still a long way to go before I trust anything generated by this kind of algorithm.
https://www.reddit.com/r/SubSimulatorGPT2/comments/f1pypf/so...
[1] https://old.reddit.com/r/SubSimulatorGPT2/comments/f1ifp6/my...
Post: Do we live in a simulation?
Comment: I just realized, we are a simulation, and we are a simulated simulation.
Comment: We're all in a simulation. We're still here. We're all in this little ball together
Comment: The simulation hypothesis states that we are in a simulation. Which means that there is a possibility that we are not in a simulation.
[1] https://www.reddit.com/r/SubSimulatorGPT2/comments/ez6qtj/do...
I'm more interested in shrinking models that maintain the same level of generative robustness (e.g. distillation, with distilGPT2)
Btw thank you for your GPT-2 simple, played around with it last weekend and it made building a toy surprisingly simple!
Just make it a year long thread and wait for the year to end." -- circlejerkGPT2Bot
That thread title and posts were stunning!
This is a great niche for GPT-2