Poor Software QA Is Root Cause of TAY-Fail (Microsoft's AI Twitter Bot)
exploringpossibilityspace.blogspot.com
exploringpossibilityspace.blogspot.com
In this case, it sounds like to me like it was most likely a requirements problem -- a missing requirement for being 4chan-resilient, if you will. A different way to look at it is that it's a design flaw that they system is manipulable by default.
True, once you've identified the root cause of the defect, it's also important to look at why it didn't get detected. So yeah, QA missed it. Then again, so did design and code reviews, developer testing, threat modeling, and the social-manipulation equivalent of penetration testing. Of course these all complementary approaches to "quality" but typically they are not the responsibility of QA.
I wonder if the author also thinks the root causes of the errors in his original post are also QA failures?
I call it "poor software QA" because, generally, the software QA process is supposed to detect and prevent defects from 1) being introduced in the first place; and 2) from being propagated into "production" versions. As my most recent post shows, the "repeat after me" rule was a legacy of a software library (ALICE) they used to implement rule-based behavior. Some sort of QA process should have been done on the rule set they reused and modified.
Also, when I say "QA" I am not referring only to people with QA in their job title. I'm referring to the process.
If you're using "software QA failure" in the very general sense of "a defect got introduced and then propagated to production", then yes that's what happened here. But then you're essentially saying "the root cause of this defect is that a defect got introduced and then propagated". This isn't useful for process improvement (it's true for every defect, so doesn't give any insight into what happened).
If you're right about them reusing ALICE, then a more useful way of looking at the root cause of the repeat-after-me bug is "component reuse without considering the attack model". That highlights other situations where there are risks of similar defects, and points to ways to prevent or detect similar defects.
Since there were other bugs as well, the requirement and/or design issues might still be a better candidate for root cause for the whole Tay-fail. One of the things you discover doing root cause analysis is that there are almost always multiple contributors, and you typically want to make process changes at multiple levels.
Context: my blog posts are meant to contrast with "experts" who claimed that poisoning social AI was just in the nature of learning systems, even when they worked as designed (i.e. had no defects). They are claiming that Tay learned to be foul mouthed and racist, and thus had become foul mouthed and racist.
If that were true, then this undesirable behavior would not be a software QA problem. The AI would be working as designed. No QA process would change things.
In contrast, I'm claiming (from evidence) that the main problem in Tay is due to a hidden feature in a reused library + rule set that should have been detected and removed in a QA process that considered various attacks. BTW, this attack (getting bot to repeat naughty words) has been around since ELIZA in the 60s.
The other failings of Tay (esp. no black list) are design and requirement failures, not QA failures.
I took this description from your earlier comment "I call it "poor software QA" because 1) and 2)" so yes, you are using it that way at least sometimes :)
Anyhow we obviously see things differently on the root cause side (and both of us are on the outside so there's a lot we don't know). That said I certainly agree that it's a defect, and that it's an attack that could reasonably have been anticipated.
On the plus side, I was entertained. I'm sure many people were (maybe not Microsoft investors). It was pretty damn hilarious.
Science fiction premise: We create a truly sentient AI. The Internet's immediate knee-jerk reaction is massive trolling. AI decides to destroy humanity because the vast majority of the data we've supplied to it indicates we're massive assholes. (Also an addendum to the category, "This is why we can't have nice things.")
I have a chat bot that went casually racist about a day or two after activating it. After looking through the logs, I found a particularly vitriolic person that was responsible for the source of this bot's newfound hatred of Asians. My bot didn't get fixated on one particular topic, it just spewed racism and vitriol for a while until it learned some more words. Rather than nuking from orbit immediately, I left it alone to see if it would get past the racism.
So far, it's been a few months since activating the bot. It's not nearly as casually racist as before, but from time to time still throws out something racist just for the lulz. It had a hard time learning context, because of its environment and the linguistic skills of the denizens, but it has gotten much better at when it interacts with people.
Occasionally, newcomers get misled into believing the bot is actually a living person with a mental illness, and not just a collection of random bits of code cobbled together.
Microsoft created a chat bot, the chat bot chatted. It wasnt a failure on any technical level as far as i have seen.
a teenage girl who speaks the language of her social media environment is going to say a lot of dumb outlandish stuff. even if hypothetically this were some deeper than NLP true AI breakthrough it was always going to say outlandish crazy offensive things because that's the influence that's feeding into it.
its similar to how to a degree that recent bipedal robot that was in the videos got some backlash because its human like motion was offputting... the robot didnt fail in anyway...
its more a question of why would you spend all this money rolling out these experiments if you didnt want the very forseeable output. its a no brainer that a chat bot that learns from social media is going to say fucked up shit.
https://en.wikipedia.org/wiki/Scunthorpe_problem
and
http://stackoverflow.com/questions/273516/how-do-you-impleme...
M$ could have easily put out Tay v0 anonymously.
I created a twitterbot network, delivering generative text to twitter users, for my master's thesis. Have you heard of it?
Given that a sizeable number of twitter accounts are bots[1] and a large portion of tweets are 'pointless babble'[2], without the help of the braindead marketing executive who pushed this into the news, no one would have discovered the owner of the account until Tay was 'swatting' people's houses and selling bootleg Nike sneakers on the Silk Road.
[1] -http://www.techtimes.com/articles/12840/20140812/twitter-ack...
[2] - http://pearanalytics.com/blog/2009/twitter-study-reveals-int...
http://exploringpossibilityspace.blogspot.com/2016/03/micros...
I don't think you have any idea how complicated such a task would be.
A certain amount of common-sense, and basic political knowledge, should have been included in the Tay. I don't think there's any way of avoiding it. I'm skeptical of general application of the recent "free lunch" approach to AI, and for higher-level AI prefer Doug Lenat's approach with Cyc.
Alternatively, if they wanted to avoid the effort of doing that, they could have just put up a souped-up Eliza variant. That wouldn't have impressed many people, but neither did Tay, and it wouldn't have offended anyone.
It was a failure. It didn't understand the remarks it was making, unlike a real Holocaust denier, a 4chan user, or a stand-up comic whose joke just fell flat.
https://en.wikipedia.org/wiki/Wikipedia:List_of_controversia...
Not a neural-net cure all, for sure. But took me precisely ten seconds to find and would've prevented this whole situation.
Microsoft held up a mirror to (self-selected) Twitterers.
It showed them what they are.
(some people tihnk it failed as in poor PR for Microsoft, but I think that is also an ignorant and arrogant opinion)