Watson crushes the competition in second round of 'Jeopardy'
news.yahoo.com
news.yahoo.com
"First, the category names on Jeopardy! are tricky. The answers often do not exactly fit the category. Watson, in his training phase, learned that categories only weakly suggest the kind of answer that is expected, and, therefore, the machine downgrades their significance. The way the language was parsed provided an advantage for the humans and a disadvantage for Watson, as well. “What US city” wasn’t in the question. If it had been, Watson would have given US cities much more weight as it searched for the answer. Adding to the confusion for Watson, there are cities named Toronto in the United States and the Toronto in Canada has an American League baseball team. It probably picked up those facts from the written material it has digested. Also, the machine didn’t find much evidence to connect either city’s airport to World War II. (Chicago was a very close second on Watson’s list of possible answers.) So this is just one of those situations that’s a snap for a reasonably knowledgeable human but a true brain teaser for the machine."
http://asmarterplanet.com/blog/2011/02/watson-on-jeopardy-da...
If it's just trained-up on statistical correlations between trigger phrases and likely answers in the constrained Jeopardy domain, then 90 32-core/512GB RAM servers seem like overkill.
What's remarkable and important to not take lightly is the result that it's possible to generate answers to often vague and indirect clues without understanding. That likely means that it will be possible to build useful systems for automating research and the synthesis of large amounts of data without needing to build artificial human-level intelligence.
I don't own a TV and plan on watching the Jeopardy match later online, so I'm just going to guess about Watson's performance. I think that humans abuse discovered patterns and structure in language and meaning to search through possible interpretations very quickly. Watson on the other hand uses far less structure and a room full of 200 cores to search through everything is knows much less efficiently. I feel like Watson's "strange" answers probably aren't nearly so strange when you realize it's simply being more fair to any possible answer than a human would.
What's scart is this sort of thing---a willingness to consider out of context answers---sounds pretty similar to the kind of behaviors we humans praise as creative!
Right, but does that structure really represent a "deeper" understanding or just vast and meticulous optimizations of statistical algorithms similar to Watson's? Or is there a difference?
We feel like we know how we think, but we can't actually explain it in enough detail to reproduce. Humans have a bad history of rationalization and tunnel vision. And now we discover that all the "wrong" ways to think deeply are actually the right ways to make a working AI.
If the AI can fool us into believing that it "understands" then maybe we can fool ourselves in the same way.
Of course, the implementations we build will always be vastly different from their appearance in the brain since the architectures are so extraordinarily different!
But the fact that bulk correlation mining can answer even some 'vague and indirect' questions isn't that remarkable. Jeopardy clues are a very constrained domain: short clues in English with some distinctive idioms, and short answers that are drawn from some well-defined and constantly recurring classes.
With true natural language understanding, offline searchable copies of Wikipedia and Wiktionary – 64GB, tops? – could be used to answer almost every question. Instead Watson uses 15TB of RAM and 2880 cores.
The best case I can make for Watson is that perhaps the alternatives shown are each actually the top option based on totally different understandings of the question. So in fact many other plausible answers are folded up just behind its right answer. The shown alternatives are meant to be: if the question means something else entirely.
¹ eg Wolf Blitzer: http://www.youtube.com/watch?v=DVC28oemocA
So today, we learned that machines can push buttons faster than people, and search is a great way to find answers for trivia questions. I doubt the former is a surprise to anybody alive in the past 50 years; the latter shouldn't surprise anybody who's ever used Google.
I arrived at the 6x approximation by googling around for avg. ethernet latencies. I'm consistently seeing numbers of .3 - .35 ms for an ethernet ping/pong. I think it's fair to assume that with the money IBM has invested in this, Watson is on at least ethernet quality connections.
Without some way to account for the fact that the human nervous system CANNOT beat Watson to the buzzer on any sort of consistent basis, the game is far less compelling than it should be.
I'm an IBMer and I think Watson has been extremely impressive, but as a Jeopardy fan it gets tiresome to see Watson win the race to the buzzer this often.
Finally, I'm hoping that tonight features more of the wordplay-heavy clues that I hear were present on the first night (why oh why was that on Valentine's Day???), because that is the element that excited me most when I heard about the challenge.
I'm not gonna lie... it was pretty entertaining watching Jennings squirm every time Watson beat him to the buzzer.
Also, there wasn't any mention of Watson having adaptive artificial intelligence but I would guess it's safe to say IBM was smart enough to include something like that. That in itself would be crazy hard to implement given the magnitude of what it's already doing... but not impossible. Maybe there are a few corrective algorithms in there somewhere.
That essay focuses on social science, but I think it's still relevant here. Experts in the area did not think they could do this, and even the people involved weren't sure. It's easy to dismiss this as "yeah, it's just a big search engine" once you already know it's been done. Besides not accurately characterizing the approach Watson takes, that sentiment misses the fact that this was an open question.
Paradoxically, people would probably be more impressed if Watson did worse and the game was more competitive. It's like watching an NBA team play a high school team. The NBA team is so good that it looks easy despite the fact that they're that good because of decades of practice.
(Disclaimer: I work at IBM Research, and have associated biases.)
Also, if there is anyone who thinks silicon valley has the smartest people around, this type of stuff should change your mind. Facebook is short trousers compared to this. and it's just a tech demo.
Watson is an impressive achievement, but there are quite a few companies in Silicon Valley whose engineers could pull this off. It's more a matter of how much money management feels like throwing at it. It's great publicity for IBM, which has to put in a lot more effort than most Silicon Valley companies in order to look cool, but can afford it.
There's driven and then there's Smart.
Here are the pictures: http://www.wired.com/epicenter/2011/02/watson-jeopardy/?pid=...
The real challenge behind Watson is the natural language parsing. Instead of abstracting information away from their sources(like a graph), sources seem to have been left intact in sentences in Watson's memory. Watson would read through this information in a way alike to how it interprets a question, and it would try to create links and possible answers based on connections in sentences from many sources(this gives thought on why pun questions are difficult for Watson). I can't speak on behalf of the mathematical implementation of the answer choices, but this is the high level way that Watson finds answers. Those videos talk about the cool stuff behind the algorithmic challenges of Watson.
So who wants to build a real Q/A site based on this? Call it hal-18000.
You'd have to learn it to deal with thick accents like this one: http://www.youtube.com/watch?v=5FFRoYhTJQQ . Honestly, I don't know if that's possible, no matter how much training you'd put into the machine.
This is what I don't get, why should be "language processing" tied to written text? Part of the answer I know, because it's easier for computers to parse, but other than that it doesn't make sense.
Seems like Watson was able to ring in (clicker) much quicker than Ken or Brad. Any unfair advantage?
That doesn't sound like a great improvement, imho.
Also, this would be pretty artificial (no pun intended), but they could analyze previous all-human Jeopardy! episodes and figure out average buzzer response time, and perhaps incorporate that into Watson.
If Watson has confidence at the moment ringing-in is allowed, it seems it will always win. So how much time it has to achieve that confidence, as a function of how the text is fed, may be more important than buzzer mechanics.
Based on the same general idea of a common starting point that motivates waiting until after the entire question has been read before allowing any ringing in, I could see there being a tiny 'common period' where all buzzes are considered as coming in simultaneously, with the person chosen to answer then being chosen in round-robin fashion (or for maximal drama, favoring whoever is behind). After the tiny common period, it would be strictly based on first-to-buzz.
It would take away a twitch-timing factor that has been important for human champions, too, but offer more fairness with regard to computers and even people with slight ticks or timing problems.
While watching, I was actually hoping for really short questions which would cut down on Watson's time to process and possibly put Watson on more even footing with Jennings and Rutter.
It's less impressive, just like a computer sorting 1000 integers faster than a human is less impressive.
edit: added "not" after the first word
I'm afraid I don't really understand the decision to trivialize the fact that we now have a computer that can answer general-knowledge natural-language queries quickly and about as accurately as a clever person. That's a Big Deal.
None of this really minimizes IBM's accomplishment, but it absolutely means this specific presentation (Jeopardy!) lacks weight for those of us who understand what the game dynamics of Jeopardy! are. This is nowhere near as impressive a "man vs. machine" victory as was Deep Blue vs. Kasparov.
Just like flying through the sky in a metal tube, talking to someone on a different continent in realtime, or converting lightning into high-energy photons and cooking your food with it.
I really enjoyed this article from Garry Kasparov in the New York Review of Books. Spoiler: it's [partially] about how the best chess player in the world is a really good human paired with a really good computer. http://www.nybooks.com/articles/archives/2010/feb/11/the-che...
So let's all join hands with the machines and sing Kumbayah ... all watched over by machines of loving grace... of course, yes, this is a medium-term view. The long-term, I suppose, probably belongs to the machines.
I mean, if they had changed the ring-in system to work in a manner befitting a man-versus-machine match, you're telling me that wouldn't be entertaining? I would be fascinated by that match. This one was a letdown.
A good analogy, similar to your car vs. human analogy, is this: Hold a competition where two contestants much first pass some sort of Turing-like test. Upon passing that test, both contestants must sort one million 32-bit integers. The first contestant to finish both tasks wins.
In that hypothetical competition, it would certainly be impressive for a computer contestant to pass the Turing-like test. (The human contestant would probably have no trouble doing so.) But the sorting task would seal the door, and the computer would win every time it is able to pass the Turing-like test. The sorting task is known to be dominated by computers, just like buzzer reflexes are known to be dominated by computers.
84% is, to my mind, very high. I’m skeptical of claims that even the best humans are better than that.
Watson is being granted first crack at the questions 90% of the time because of its electromechanical advantage. IBM may not have the mean brainpower that Google has, but they can clearly build a computer that can press a button quicker than Ken Jennings.
Knowing that a computer can consistently beat even the best to ever play the game to the buzzer, the IBM team could be pretty well assured of success once they got Watson performing well enough.
"IBM holds more patents than any other U.S.-based technology company and has nine research laboratories worldwide. Its employees have garnered five Nobel Prizes, four Turing Awards, nine National Medals of Technology, and five National Medals of Science."
admittedly, they've been around longer, but they're not exactly playing with crayons over there.
But instant button pressing is not terribly impressive and yet is Watson's critical advantage. This match is framed as a battle of knowledge, not a struggle for humans to overcome the massive disadvantage of their meat-based nervous system, which it is.
This is largely glossed over. Watson would lose horribly if one of its engineers was pressing the button. If its electronic advantage on the button was taken away, it would be the kind of fascinating match I and a lot of people were expecting.
"what is the capital of china"
which are probably closest to Jeopardy questions (technically speaking.)
The team that does this is a lot more than seven Google engineers. Yet, it is not the top selling features of Google, because -- it's not that easy.
There is a fine line between "seems easy" and "nearly impossible". A Google engineer does not equate to infinite skills.
I just hope the next challenge isn't "defeat a standing army"
That shouldn't minimize the accomplishment of actually performing well enough. If Watson were fast on the buzzer but couldn't answer accurately, he's be pretty far in the red about now.
Therefore they've made a machine that plays Jeopardy about as well as an average Jeopardy contestant.
That's still an impressive feat but not quite one I would have thought was out of reach five years ago.
They have a ton of priors, a ton of specific optimizations, and a lot of resources.
It's pretty impressive but in a they-clearly-worked-a-lot-on-this way, not a what's-their-secret kind of way, you know?
If this isn't a "what's their secret kind of way" then what is, in your opinion? Or are you just more of the pragmatic, if its been done then obviously it could be done, duh.
http://www.madpickles.org/rokjoo/2011/02/14/ibm-watson-vs-go...
And let me add, if this is so incremental, I'd love to Google just turn this feature on. Let me ask it any question and just have it return the answer. at the top of the results page w/o me having to click a link. I don't think we'll be seeing that from Google any time soon.
During tournaments of champions, Jeopardy! is not about how many correct responses you can come up with, as all competitors will know the vast majority of them. It's all about timing. I don't know the exact statistics, but it seemed like Watson knew about 75% of the correct responses. I strongly suspect the two human contestants knew a greater percentage. On day 2, it just came down to timing. Watson was only beaten to the buzzer three times when it knew the correct response.
It didn’t know the correct response or was not confident enough to answer 16% of the time (5 of 31). Nobody at all could give an answer to 6% of all questions (2 of 31).
If the two questions we know nobody could answer are excluded, Watson knew the right answer 90% of the time (26 of 29).
Do individual Jeopardy contestants correctly know more than 84% of the answers? Two questions were left unanswered correctly by everyone, 93% (29 of 31) is already the upper bound for the human performance in this round, that’s not very far from 84%.
Looking at the six rounds of Jeopardy the J! Archive has of games with Ken Jennings and Brad Rutter [0] we see that the average upper bound in those six rounds was 87% or 26 questions (of 30). The minimum was 23 and the maximum 28 (of 30). That’s, again, very much an upper bound.
Looking at all this, I’m pretty confident that Watson would be doing well even if it had human reaction times. (This is how I would change the game: Within 200ms or so – whatever the human reaction time on a Jeopardy buzzer is — after the buzzers are open, the player to answer is randomly selected from all who managed to buzz in. After those 200ms everything stays the same. Players are also not punished for buzzing in too early. This is a minimal change to the game, not a completely different game which I think is important.)
And making it harder (more obscure trivia) would only help Watson.
Whether or not we can make a more interesting or "fair" game of Jeopardy involving such a machine is an entirely separate, and to my mind far less interesting, question.
Where do we go from here? It would seem silly to say "we make a more interesting Humans vs. machines Jeopardy game." Rather, it seems more prudent to figure out ways to expand on this research and use the underlying technology to solve more interesting, and practical, problems.
Place all three contestants in isolation from each other.
All three hear the question read, and buzz-in just as they do now.
Allow ALL contestants who buzz in to answer the question, but do not allow them to know about their opponents' performances.
Record all contestants' buzz-in reaction times.
At the end of the game, compare only the accuracy of answers to determine the winner.
At the end of the game, compare buzz-in reaction times to see how thumbs fare against relays.
https://spreadsheets1.google.com/ccc?key=tth_jhM8vyBAuogqHll...
Popular shows are reshown on weekends. Also, around holidays and other non-normal weeks of broadcasting, older shows are re-aired.
I don't believe episodes are available legally online.
Mubarak's not president of Egypt anymore!