T5: The Text-to-Text Transfer Transformer
ai.googleblog.com
ai.googleblog.com
Options:
1. "Mechanical Turk" style, a massive undertaking to manually clean up Common Crawl, perhaps using underpaid labor in third world countries (such as samasource.com does)
2. By means of somehow getting the internet to do it for them with something like reCAPTCHA
3. With the help of machine learning / traditional text processing
4. Some other way
Anyone has any ideas? I'm intrigued. The paper [https://arxiv.org/pdf/1910.10683.pdf] and the website [https://www.tensorflow.org/datasets/catalog/c4] mention almost nothing, except for an option to switch off the cleaning & deduplication, which hints at option number 3.
> Unfortunately, the majority of [the text in Common Crawl] is not natural language. Instead, it largely comprises gibberish or boiler-plate text like menus, error messages, or duplicate text. Furthermore, a good deal of the scraped text contains content that is unlikely to be helpful for any of the tasks we consider (offensive language, placeholder text, source code, etc.). To address these issues, we used the following heuristics for cleaning up Common Crawl’s web extracted text:
•We only retained lines that ended in a terminal punctuation mark (i.e. a period, exclamation mark, question mark, or end quotation mark).
•We removed any page that contained any word on the “List of Dirty, Naughty, Obscene or Otherwise Bad Words”. [https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and...]
•Many of the scraped pages contained warnings stating that Javascript should be enabled so we removed any line with the word Javascript.
•Some pages had placeholder “lorem ipsum” text; we removed any page where the phrase “lorem ipsum” appeared.
•Some pages inadvertently contained code. Since the curly bracket “{” appears in many programming languages (such as Javascript, widely used on the web) but not in natural text,we removed any pages that contained a curly bracket.
•To deduplicate the dataset, we discarded all but one of any three-sentence span occurring more than once in the dataset.
Additionally, since most of our downstream tasks are focused on English-language text, we used langdetect [https://pypi.org/project/langdetect/] to filter out any pages that were not classified as English with a probability of at least 0.99.
Looking at that list, I wonder what the unintended consequences of a decision like this is. If you want to create something related to sentiment analysis, that swear words you discarded is a useful signal, not noise right? If you wanted to use the dataset somehow for your tour guide business in Austria, how does it handle the the village called Fucking? Does T5 understand the British colloquialism for cigarettes? Can ornithologists talk to it about penguins and eagles, but not about yellow-bellied tits and blue-footed boobies?
We are also talking to Common Crawl to see if they will host a prepared copy since we do not have redistribution rights.
Trivia fact recall and NLP seem like two quite different tasks even though both are required to do well in a quiz.
> Q: How did Gus Grissom, Ed White and Roger B. Chaffee die in 1967?
> You: "Apollo 1" WRONG
> T5: "They were killed when their Apollo 1 spacecraft exploded" WRONG
> Correct answer: burned to death
> Q: Which Alpine peak is known in Italy as Monte Cervino?
> You: "Monte Cervino" CORRECT
I wonder how many of the problems with this game could be fixed by applying T5 itself to the answer grading.
More likely: suffocated (as with most fire deaths) https://history.nasa.gov/Apollo204/invest.html:
”d. MEDICAL ANALYSIS
Loss of consciousness was due to cerebral hypoxia due to cardiac arrest resulting from myocardial hypoxia. Factors of temperature, pressure and environmental concentrations of carbon monoxide, carbon dioxide, oxygen and pulmonary irritants were changing extremely rapidly. It is impossible to integrate these variables on the basis of available information with the dynamic physiological and metabolic conditions they produced, in order to arrive at a precise statement of time when consciousness was lost and when death supervened. The combined effect of these environmental factors dramatically increased the lethal effect of any factor by itself. It is estimated that consciousness was lost between 15 and 30 seconds after the first suit failed. Chances of resuscitation decreased rapidly thereafter and were irrevocably lost within 4 minutes.”
Imho a bit unfortunate, as is calling the decoder or the encoder of a transformer "a transformer", as it has happened with GPT and BERT, which now forces people to use "full transformer" or using phrases like the title of the blog post.
Watson (Jeopardy Watson, not the IBM branding exercise Watson is now) has much weaker text understanding models, but has much much better optimisations for the incremental style of data release that you see in Jeopardy (ie, you get more and more data the longer you listen). IBM did a lot of work optimising when to answer as well as trying to get the correct answer.
The closest analogy that is regularly studied in modern QA research is "Quizbowl"-style datasets, but these tend to be much smaller than the SQUAD datasets that most modern neural network QA systems are built against.
"summarize: state authorities dispatched emergency crews tuesday to survey the damage after an onslaught of severe weather in mississippi..."
And the output is:
"six people hospitalized after a storm in attala county."
That's quite a bad summary, no mention of "six" people in the original text, no mention of hospitalization. And "attala" county is too specific, a precision not present in the original text.
If that's the result of their model, that's not good. If it's coming from the training set, it's an even bigger problem. I guess it's the result of the model, because some issues can be explained by correlations ("emergency" correlates with "hospital", "mississipi" correlates with "attala").
I'm wondering why they chose this example for the flagship figure of their paper.
I'm not sure if the information on http://nlpprogress.com/english/machine_translation.html is accurate, but it appears that the top translation results rely on backtranslation, boosting and other data augmentation techniques with a vanilla transformer model. It would be interesting to see the Bleu scores for T5 that's more optimized for translation specifically.
> Colossal Clean Crawled Corpus (C4)
pun detector at 3.6 punits. not great, not terrible.
> Q: What is the opposite of an acid?
> You: alkaline
> T5: Alkali
> Correct answer: a Base
:-)
If you need to tweak 11-billion parameters to get a particular result, I don’t see how you can call whatever is being called a model, more like a component of a model.
So exactly like the "unix pipe" philosophy invented 47 years ago?
I guess ideas are cyclical...