Natural language generation: The commercial state of the art in 2020
cambridge.org
cambridge.org
He give a list of commercial providers and concludes that most of them just offer smart templates. This is the important part:
> To the extent that you can tell from the clues to functionality that are surfaced by these various products, all the tools are ultimately very similar in terms of how they work, which might be referred to as ‘smart template’ mechanisms. There is a recognition that, at least for the kinds of use cases we see today, much of the text in any given output can be predetermined and provided as boilerplate, with gaps to be filled dynamically based on per-record variations in the underlying data source. Add conditional inclusion of text components and maybe some kind of looping control construct, and the resulting NLG toolkit, as in the case of humans and chimpanzees, shares 99% of its DNA with the legal document automation and assembly tools of the 1990s, like HotDocs (https://www.hotdocs.com). As far as I can tell, linguistic knowledge, and other refined ingredients of the NLG systems built in research laboratories, is sparse and generally limited to morphology for number agreement (one stock dropped in value vs. three stocks dropped in value).
That sounds pretty negative, but he emphasizes that putting an easy-to-use UI on well understood technology is meeting real business needs.
At the end he briefly touches on GPT2 and related technology.
Generation is easier than understanding. Systems like GPT-3 are not capable of respecting constraints, which is this basis for practical creativity.
I provide raw data for some news services and that data could be shown as tables or simple dashboards that are faster to read for users, but my customers insist in generating as much text as possible from, sometimes, 2 values.
The incentives are totally misaligned, news services try to get as much time from it's readers as possible. It's quite depressing.
[1] "Quality" here meaning how it looks to Google's crawler, not actual quality content.
In particular at search time, google has a vested interest in limiting both the number of links to a particular domain and the occurrence of content farm links when Wikipedia would suffice. Content farms are pretty detectable in any objective relevance annotation workflow.
Good data-to-text NLG applications not only summarise data but they also can provides insight into causal relations of why events occurred by leveraging domain knowledge.
I think NLG is really interesting, the problem is that there are incentives to create long form content that doesn't add any value to readers.
I am eager to see if neural approaches develop to better handle constraints, but part of me thinks the level of control over generated prose that template based approaches provide is essential for customer-facing text and that part of the market will persist even when every other recipe in Google search results is GPT-3 generated rubbish around the actual damn list of ingredients.
https://github.com/google-research/text-to-text-transfer-tra...
I must completely lack imagination, because I don’t know what to use it for if it doesn’t give me access to the weights.
Apparently it was published just one day before GPT-3.
> Sometimes the results produced by GPT-2 and its ilk are quite startling in their apparent authenticity. More often than not they are just a bit off. And sometimes they are just gibberish. As is widely acknowledged, neural text generation as it stands today has a significant problem: driven as it is by information that is ultimately about language use, rather than directly about the real world, it roams untethered to the truth. While the output of such a process might be good enough for the presidential teleprompter, it would not cut it if you want the hard facts about how your pension fund is performing. So, at least for the time being, nobody who develops commercial applications of NLG technology is going to rely on this particular form of that technology.
But I agree there are some applications it is useful for, like education.
I don't know why that would imply that it knows nothing about the real world, unless the data corpus it is trained on likewise bears no relation to reality...
It’s trained on Reddit, so I wouldn’t rule that out.
As an experiment I used a GPT-3-powered website [1] to see what GPT-3 has to say about bears, and the first answer was:
> "Weird that every day, there are so many cute/funny/entertaining bears to enjoy online but hardly any on the ground."
When asked about beards, the first answer has no relation with beards at all:
> "If a person doesn’t constantly outwit, outplay, outlast, others, the strong eat the weak."
And then there's that time when GPT-3 told someone to kill themselves [2].
While funny and (mostly) grammatically correct, these "thoughts" are nonsense and no amount of extra parameters is going to solve the disconnection between GPT-3 and reality. I imagine you could condition GPT-3 to generate text for a specific piece of data in such a way that guarantees the correctness of its output, but at that point you might as well throw GPT-3 away and write a rule-based system.
The site you tried is a tweet generator, not a question answering site. I prompted GPT-3 with "Bears and beards are different because" and got...
"Bears and beards are different because they are not the same thing.
Bears are animals. Beards are facial hair.
Bears are dangerous. Beards are not.
Bears live in the woods. Beards live on your face.
Bears eat people. Beards do not."
But my original point was mainly that this field is moving fast and the the old school NLG companies (I created one back in the day!) are toast.
Yes it is true that bears are animals.
No it is not true that, as GPT-3 said, "There aren't any on the ground"
It's younger sibling DALL-E is capable of language grounded in images, I expect the next version to be multi-modal as well. On another line of research there's effort to tame the horse (GPT) by attaching a secondary neural net. This can monitor language, topic, style and bias and ensure increased accuracy in tasks by auto-learning good prompts. It would make development of applications much easier because the base model which was super expensive to train can be reused many times while the secondary net is small and fast to train. Other efforts are related to including a search engine on an inner loop, to make the language model able to query large collections. Also, there's an open effort to create a huge text corpus, so far 800GB (The Pile). It improves on the GPT-3 training corpus on some categories that were lacking.
I think it's safe to say the article is way off the current research level.