Paper: https://arxiv.org/pdf/2004.13637.pdf
Open Source: https://parl.ai/projects/recipes/
Ask us anything, the Facebook team behind it is happy to answer questions here.
Paper: https://arxiv.org/pdf/2004.13637.pdf
Open Source: https://parl.ai/projects/recipes/
Ask us anything, the Facebook team behind it is happy to answer questions here.
I have a few questions: Could this be used as a tool to get a feel for public sentiment? For example, could you ask the bot what it thinks about gun control and have it spit out a policy that appeals to the common public? If you ask the bot what it thinks about how a company will perform, how accurately does it predict? I know that the model will contain the biases of the data set, but I'm curious if you've run these types of experiments. What do you think the results would be if you had an even bigger, more diverse corpus? (devil's advocate, for the sake of discussion: perhaps everyone's fb messenger and WhatsApp chat history)
Finally, you have clearly gone to great lengths to make the bot pleasant to interact with. What sort of results to you get when you train such a huge model on an uncurated corpus and don't try to tweak its personality? I find myself wishing that you didn't try to do this as the bot seems to be hyper-agreeable. I. E Too many responses like "You like watching paint dry? That's super interesting! I love watching paint dry!".
In the paper, we did explore what happens when you do NOT fine tune it on the specialized tasks (knowledge, empathy and personality). The non-finetuned bot was both less engaging and more toxic. The special finetuning is really important to getting this bot to be as high quality as it is.
It's just a matter of time before a model of this size can be run on commodity hardware and somebody will take the brakes off and/or attempt to run experiments that aren't just "can this thing pass the turing test?". I'd be really interested to know the thoughts of the team, given their expert knowledge and experience with the matter.
Unfortunately you can't talk to it. (I've wanted to retrain a version that you can interact with dynamically, someday.)
plus motherboard, cpu, ram etc.
Why do you think that Facebook paid for this research?
I mean, I don't see a plain facebook connection, but I see googleads and amazon connection and don't know what else obscure things. I doubt it would be hard to sneak something in, that just checks whether this is the same browser where you just switched over from the facebook tab (where the onblur event just got captured). But again, I am not an expert in tracking ads, nor reddit, I just know the web and its various data transmitting technices quite good.
This seems like wishful thinking. Having knowledge and having the resources to do something with it are two very different things.
I fail to see where playing out the "but others are doing it too" card exempts the responsibility of those who either lower or eliminate the barrier to entry to these attacks.
every AI sound/word/picture editor i've ran into says something along the lines of "we're releasing this data set to help stay secure in this day and age of easy counterfeiting of X.", but they never really mention how you apply the data in an adversarial way against itself -- they just sort of hand-wave that part.
Same with fake AI generated Obama video and sound, and earlier data-set generated chatbots; it's plastered all over the projects things like "Since these methods are available we think that it's important that this data is disseminated so that other's can use it to validate real world data sources", but again -- how?
We have the real data, we have the fake data -- how is this diff done, exactly?
I'm willing to bet it isn't as easy as all the AI researchers who release this stuff claim it may be.
With the data public its more akin to driveby ssh login attempts. Not being important doesn’t mean your not under attack and people can take the necessary precautions.
There are few reasonable ways to "take precautions" against nuclear weapons and there are few reasonable ways to "take precautions" against something like this short of swearing off of social media entirely.
Without reasonable defences, all you really accomplish is ramping up proliferation.
I agree with your approach though.
Just use the bot. Put it into action. Let the bot answer HN's questions
Still, unlike Tay, we purposely did not create a service for it and do not advise creating one. This is for research purpose only and more effort needs to be made on safety before it can me more broadly consumed.
As far as facts go, we also fine-tuned the model on the Wizard of Wikipedia task (https://arxiv.org/abs/1811.01241), which helps improve its knowledgeability. Again, it still isn't perfect.
It's way too heavy/expensive to host as an individual.
https://en.m.wikipedia.org/wiki/Tay_(bot)
(but I am not sure if this bot does indeed learn from new conversations)
Why do you choose to use your incredible talents for the benefit of such a disgusting company?