Seems like Watson was able to ring in (clicker) much quicker than Ken or Brad. Any unfair advantage?
Seems like Watson was able to ring in (clicker) much quicker than Ken or Brad. Any unfair advantage?
During tournaments of champions, Jeopardy! is not about how many correct responses you can come up with, as all competitors will know the vast majority of them. It's all about timing. I don't know the exact statistics, but it seemed like Watson knew about 75% of the correct responses. I strongly suspect the two human contestants knew a greater percentage. On day 2, it just came down to timing. Watson was only beaten to the buzzer three times when it knew the correct response.
It didn’t know the correct response or was not confident enough to answer 16% of the time (5 of 31). Nobody at all could give an answer to 6% of all questions (2 of 31).
If the two questions we know nobody could answer are excluded, Watson knew the right answer 90% of the time (26 of 29).
Do individual Jeopardy contestants correctly know more than 84% of the answers? Two questions were left unanswered correctly by everyone, 93% (29 of 31) is already the upper bound for the human performance in this round, that’s not very far from 84%.
Looking at the six rounds of Jeopardy the J! Archive has of games with Ken Jennings and Brad Rutter [0] we see that the average upper bound in those six rounds was 87% or 26 questions (of 30). The minimum was 23 and the maximum 28 (of 30). That’s, again, very much an upper bound.
Looking at all this, I’m pretty confident that Watson would be doing well even if it had human reaction times. (This is how I would change the game: Within 200ms or so – whatever the human reaction time on a Jeopardy buzzer is — after the buzzers are open, the player to answer is randomly selected from all who managed to buzz in. After those 200ms everything stays the same. Players are also not punished for buzzing in too early. This is a minimal change to the game, not a completely different game which I think is important.)
And making it harder (more obscure trivia) would only help Watson.
I mean, if they had changed the ring-in system to work in a manner befitting a man-versus-machine match, you're telling me that wouldn't be entertaining? I would be fascinated by that match. This one was a letdown.
It's less impressive, just like a computer sorting 1000 integers faster than a human is less impressive.
edit: added "not" after the first word
I'm afraid I don't really understand the decision to trivialize the fact that we now have a computer that can answer general-knowledge natural-language queries quickly and about as accurately as a clever person. That's a Big Deal.
None of this really minimizes IBM's accomplishment, but it absolutely means this specific presentation (Jeopardy!) lacks weight for those of us who understand what the game dynamics of Jeopardy! are. This is nowhere near as impressive a "man vs. machine" victory as was Deep Blue vs. Kasparov.
Just like flying through the sky in a metal tube, talking to someone on a different continent in realtime, or converting lightning into high-energy photons and cooking your food with it.
I really enjoyed this article from Garry Kasparov in the New York Review of Books. Spoiler: it's [partially] about how the best chess player in the world is a really good human paired with a really good computer. http://www.nybooks.com/articles/archives/2010/feb/11/the-che...
So let's all join hands with the machines and sing Kumbayah ... all watched over by machines of loving grace... of course, yes, this is a medium-term view. The long-term, I suppose, probably belongs to the machines.
A good analogy, similar to your car vs. human analogy, is this: Hold a competition where two contestants much first pass some sort of Turing-like test. Upon passing that test, both contestants must sort one million 32-bit integers. The first contestant to finish both tasks wins.
In that hypothetical competition, it would certainly be impressive for a computer contestant to pass the Turing-like test. (The human contestant would probably have no trouble doing so.) But the sorting task would seal the door, and the computer would win every time it is able to pass the Turing-like test. The sorting task is known to be dominated by computers, just like buzzer reflexes are known to be dominated by computers.
84% is, to my mind, very high. I’m skeptical of claims that even the best humans are better than that.
Watson is being granted first crack at the questions 90% of the time because of its electromechanical advantage. IBM may not have the mean brainpower that Google has, but they can clearly build a computer that can press a button quicker than Ken Jennings.
Knowing that a computer can consistently beat even the best to ever play the game to the buzzer, the IBM team could be pretty well assured of success once they got Watson performing well enough.
"IBM holds more patents than any other U.S.-based technology company and has nine research laboratories worldwide. Its employees have garnered five Nobel Prizes, four Turing Awards, nine National Medals of Technology, and five National Medals of Science."
admittedly, they've been around longer, but they're not exactly playing with crayons over there.
I just hope the next challenge isn't "defeat a standing army"
But instant button pressing is not terribly impressive and yet is Watson's critical advantage. This match is framed as a battle of knowledge, not a struggle for humans to overcome the massive disadvantage of their meat-based nervous system, which it is.
This is largely glossed over. Watson would lose horribly if one of its engineers was pressing the button. If its electronic advantage on the button was taken away, it would be the kind of fascinating match I and a lot of people were expecting.
"what is the capital of china"
which are probably closest to Jeopardy questions (technically speaking.)
The team that does this is a lot more than seven Google engineers. Yet, it is not the top selling features of Google, because -- it's not that easy.
There is a fine line between "seems easy" and "nearly impossible". A Google engineer does not equate to infinite skills.
That shouldn't minimize the accomplishment of actually performing well enough. If Watson were fast on the buzzer but couldn't answer accurately, he's be pretty far in the red about now.
Therefore they've made a machine that plays Jeopardy about as well as an average Jeopardy contestant.
That's still an impressive feat but not quite one I would have thought was out of reach five years ago.
They have a ton of priors, a ton of specific optimizations, and a lot of resources.
It's pretty impressive but in a they-clearly-worked-a-lot-on-this way, not a what's-their-secret kind of way, you know?
If this isn't a "what's their secret kind of way" then what is, in your opinion? Or are you just more of the pragmatic, if its been done then obviously it could be done, duh.
http://www.madpickles.org/rokjoo/2011/02/14/ibm-watson-vs-go...
And let me add, if this is so incremental, I'd love to Google just turn this feature on. Let me ask it any question and just have it return the answer. at the top of the results page w/o me having to click a link. I don't think we'll be seeing that from Google any time soon.
Also, this would be pretty artificial (no pun intended), but they could analyze previous all-human Jeopardy! episodes and figure out average buzzer response time, and perhaps incorporate that into Watson.
If Watson has confidence at the moment ringing-in is allowed, it seems it will always win. So how much time it has to achieve that confidence, as a function of how the text is fed, may be more important than buzzer mechanics.
Based on the same general idea of a common starting point that motivates waiting until after the entire question has been read before allowing any ringing in, I could see there being a tiny 'common period' where all buzzes are considered as coming in simultaneously, with the person chosen to answer then being chosen in round-robin fashion (or for maximal drama, favoring whoever is behind). After the tiny common period, it would be strictly based on first-to-buzz.
It would take away a twitch-timing factor that has been important for human champions, too, but offer more fairness with regard to computers and even people with slight ticks or timing problems.
While watching, I was actually hoping for really short questions which would cut down on Watson's time to process and possibly put Watson on more even footing with Jennings and Rutter.
That doesn't sound like a great improvement, imho.
Whether or not we can make a more interesting or "fair" game of Jeopardy involving such a machine is an entirely separate, and to my mind far less interesting, question.
Where do we go from here? It would seem silly to say "we make a more interesting Humans vs. machines Jeopardy game." Rather, it seems more prudent to figure out ways to expand on this research and use the underlying technology to solve more interesting, and practical, problems.