IBM‘s Project Debater does debate club-style discussions with humans
theverge.com
theverge.com
Seriously, LOOK at this:
https://i.imgur.com/an7857N.png
I have a 1920x1080 24" display and you can't fit more than eight words in width on the screen and feel the need to hide all the navigation in a hamburger menu that disappears when scrolling? This is some of the actual worst design I've seen, ever.
There's random call-to-action buttons, spaces where I couldn't tell if something was loading or if the blank was intentional, things fading and moving in too slowly, and animations that are just plain weird.
To me it looks like bold design that got a bit too...bold.
Edit: Don't get me wrong, the Plex typeface is great, but the website dedicated to it is wonky at best.
https://www.youtube.com/watch?v=PkSzmnA1CQQ https://www.youtube.com/watch?v=ZIY1uSxL-qQ
But now I am suspicious that it always debates the same people (they make reference to having worked with the machine for over a year). If that is the case, then they know precisely what to say to trigger certain canned responses (it makes jokes that were quite obviously well-planned by a human). The rest of the material seems to just be content culled from online documents and read out by the machine.
The achievement here, if there is one, is to parse online material and decide which snippets support a given argument and which don't.
"Watching the debate, I figured the answer was that it didn’t quite get it, but I wasn’t positive. I couldn’t tell the difference between an AI not being as smart as it could be and an AI being way smarter than I’ve seen an AI be before. It was a pretty cognitively dissonant moment. Like I said, unsettling."
... reminds me of when Kasparov lost to Deep Blue, when it made a strange move. He thought the move was too sophisticated for a computer. Fifteen years later, one of Big Blue's designers said the move was the result of a bug in Deep Blue's software.
IBM is a marketing machine.
I, for one, am not excited about the prospect of AI-generated logical fallacies.
Dismantling such is where the real work is at. Give me a heads up when they automated that.
An accusation of dissembling implies intent, and that probably requires some sort of theory of mind, to make assumptions about how the response will be received. The author of the article seems to be anthropomorphizing, attributing greater cognitive powers to the system than it is displaying.
Welser also says "there’s been no effort to actually have it play tricky or dissembling games", leaving open the question of whether it has nevertheless been exposed, in training, to arguments that avoid the issue, and whether that has influenced the way in which it constructs and scores candidate replies.
If I tell a random person on the street that my AI can play checkers, they think, "A human can play checkers. A human can also drive a car. Therefore, if an AI can play checkers it must also be able to drive a car."
People use AI tasks as markers that AI is AGI / human, then extrapolate to logical abilities it must have.
Which, if you don't know what a matrix is, is fairly understandable.
This would be a dishonest/manipulative debate AI paired with big data and microtargeting. It's like pairing a trained and primed con artist with every single human being in parallel and having it shadow and work on them 24/7. Big data and cross agent "fleet wide learning" would allow it to get smarter very rapidly.
I consider this scenario to be as much of an existential threat as atomic and biological weapons.
Dismantling such is where the real work is at. Give me a heads up when they automated that.
Now all the AIs know how to earn chmod775's trust.
Pretty much everyone who works in tech knows that IBM Marketing has made much more progress in technology than IBM AI. Please don't add to the noise and confusion the average person is already under by writing an article so askew from reality. Even from the limited and carefully trimmed Debater quotes that you chose to include, it is apparent that what's going on is merely a gimmick. It is also apparent that you cherry picked the best sounding responses, then stripped any context away from them that might have exposed your charade.
So again, shame on you Dieter Bohn, for deceiving your audience. As a journalist you have a responsibility to educate and inform, and you have failed at that responsibility I'm this article.
I wonder if anybody did a sort of reverse Turing test - just like regular one but no computers involved at all, unknown to the tester which is told one of them is an AI, and how many people would be declared failing it?
> Another IBM researcher suggested that this technology could help judge fake news.
That part is actually scary. Opaque algorithms from IBM would decide for us what are facts and what are fakes? And they'd know how? Because IBM marketing dept said they're so good? Thanks but no thanks.
It's so discouraging hearing all about Watson and health for past X years and then reading AI insider articles refuting the work as total bullshit and a PR job to keep IBM relevant today.
Don't know why this would be any different.
IBM did build a computer that could play jeapordy. But being able to play jeapordy doesn't tell us how good any of the subcomponents are at, e.g. understanding doctors notes/journal articles.
The same thing goes here, they built a system that can debate with some ability. Can it be used to spot fake news as a researcher says? Maybe some subcomponents could be reused for that, but there's no guarantee that they will be good enough for that task either.
You really have to take modern AI at face value and not extrapolate from what it actually does to anything else, otherwise you will whipsaw between optimism and cynicism.
This was a demo that IBM could synthesize a plausible argument that supports or refutes a given assertion. That's interesting and a bit impressive (presuming its argument isn't simply a regurgitation of some position paper it found online). But without understanding cause and effect, its arguments will remain very superficial, probably driven by a small catalog of argument 'frames' (templates) that adaptively fill a handful of slots like [needs] [means] [goal] [conflicts] and [emotional hooks]. Using such simple recipes it can produce verbiage that plausibly sounds humanlike, but isn't actually reasoning. It's likely that the system couldn't even diagnose help desk problems using basic logic, e.g. backtracking to thise dependencies that might have caused the given outcome, thereby identifying only the possible and plausible causes.
Obviously I can not prove that they haven't developed a miracle device which can construct complete arguments and theories in real time, but if they had done, they'd've shown something more convincing.
Thing is that the media doesn't pay proportional attention and respect they do to other companies that haven't proven anything yet.
Where's Google thing that they showed last I/O where an AI would call a business to schedule an appointment? Vaporware.
Some of the insider postmortems on Watson Health that are floating around on the internet are... interesting.
It's being rolled out at right now: https://www.theverge.com/2018/12/5/18123785/google-duplex-ho...
IBMs biggest problems are still cultural, middle management and recruiting / retaining skilled developers. This is why they cannot deliver on the hype their amazing marketing teams create.
They have plenty of very, very smart people. The researchers are not IBM's problem but the management that is not capable of bringing that research to market.
It seems FB and Google have had more negative societal impacts (ie fake news, public manipulation) than IBM, Salesforce, and numerous other niche AI companies.
We all need a more nuanced perspective on all thinks than this absolutism we have embraced as a society, online.
IBM are much more open to mixing up statistical machine learning and symbolic AI, and benefit from the advantages of both. For instance, Watson used a chart parser with transitions learned using deep learning [1] to fill in frames [2] from a natural-language sentence (the challenge sentence in the Jeopardy game) and also Prolog to identify relations between elements of the parse so as to figure out the sentence's context.
Other big AI players are pretty much exclusively focused on statistical machine learning and, in particular, deep learning, and then again for very specialised tasks like image recognition, speech recognition and machine translation. IBM on the other hand is building broader-scope systems, that perform more nuanced, less strict functions; like in this example, debating, or question answering (essentially) with Watson.
So yes, in some sense, IBM is way ahead of the others in creating powerful AI systems. Just let's not confuse IBM's research with the products it tries to sell based on that research.
___________
[1] You'll have to take my word for that- I've read the Watson research papers but my pdf copies of them are now all on a hard-drive that has since crashed and I don't remember how to find them online.
[2] A cognitive data structure first proposed by Minsky in the '70s:
https://en.wikipedia.org/wiki/Frame_(artificial_intelligence...
Yeah I don't know about that. At best they are starting to incorporate deep learning into Watson. I'm guessing Watson has always been a rules engine and more specifically one designed with a lot of statistical analysis. That's far below the level of sophistication we're at with NNs.
This is IBMs page on deep learning:
https://www.ibm.com/cloud/deep-learning
While it might be nice to use whatever that tool is - it's a wrapper on top of Tensorflow/Keras or Pytorch... Wouldn't it make sense for it to be a wrapper on top of "Watson's API" whatever that is.
There's so much confusion in the market as to what Watson is, and much of it is IBM Marketing's fault. Watson is a Natural Language Processing platform. At a low level it does all the standard NLP stuff - tokenization, part of speech tagging, lemmatization, dependency parsing, co-reference resolution, stopword removal, synonym expansion, etc. Some of this is rule based, some is ML based. Building on top of that you have stuff like an Intent Classifier, Passage Retrieval with a Ranker, Named Entity extraction, Relationship extraction, and Sentiment tagging. Building on top of that, you have the actual commercial applications - chatbot/assistant, semantic search engine, knowledge graph, etc. That's what it is. It's not vaporware or smoke and mirrors, but it's obviously not AGI either.
Sounds like they are gearing up for a launch.
The problem I have with debates in popular culture is that it's often performative and participants devolve to whoever is best at dominating a discussion, or whoever is the more eloquent with not much light at the end of the tunnel.
It does mean that topics need to be apolitical in order to ensure a fair score - no debate will turn a Republican into a Democrat. But for other topics this scoring works extremely well.
I'm not saying they are, but even when it did beat Kasparov, where was the proof, code we can run ourselves and see, witnesses, just basic accountability stuff that we can make us safely believe that these are not staged?
Does anyone know anything?
1. There were only a handful of players around during those matches who could hold a candle to Kasparov. A team of the top GMs could maybe topple him, but that would require a significant conspiracy by individuals with no clear incentive to hype-up computer chess.
2) It wasn't long until you could run open source code on your home PC that could crush Kasparov.
I suspect that much of the power of Watson comes from it's data sets and preprocessing, and that technically it wouldn't look nearly as impressive as Deep Mind's work. Open sourcing its core algorithms might not give any useful insights, especially if one lacks the data sets and a powerful enough custom-built machine.
I can't even assume we all agree winning against Kasparov is proof enough. Just for argument's sake, you can buy him out, convince him that this is an amazing PR for him as well - lose one game, be double famous for life, just as you can convince a boxer to lose a match. None of this tells me 100% that Watson is AI capable of winning a chess game in the 90s or hold an intellectual debate in 2018. Again, I'm not saying it didn't happen, I'm saying burden of proof is bothering me. A good friend of mine who is a VP Cloud at IBM said when I asked last year "no one can explain Watson at ibm - it's a black box" I also wrote a few apps with tensorflow.
Kasparov, winning Go game, playing fortnite, sure.What they are suggesting here and with Sophie, without proof of any kind, is beyond me.
Is this possible?
If so, then call it Data Science or Push-button analysis.
Is this a reference that I’m not picking up on?
Deep blue beating Kasparov obviously wasn't staged because there was no human strong enough to beat Kasparov at chess. The jeopardy thing wasn't staged (but also wasn't necessarily that impressive). To this day you can download a chess engine greater than the strongest human for free online (stockfish).
This, however, is almost certainly a human-controlled PR-Stunt.
Asking for evidence is fine, but let's not conflate this with something real like chess.
Thank you. Unless they prove to us that it is not, it is.
PS. Chess game could also be staged without having to find a better player. My other answer is above.
1. Grab key points from influential books and articles to create templated responses
2. Do no original research and thus be completely unable to rebut criticisms of the source data or analysis
3. Counter #2 above by falling back on crowd-pleasing templated responses, forcing the moderator to move the debate along
4. Never, ever concede anything, because "debates" are no longer about listening to an argument and carefully considering their merits, it is a time-boxed verbal combat sport where there must be a winner and a loser.