Lee Sedol Beats AlphaGo in Game 4
gogameguru.com
gogameguru.com
Lee Sedol is playing brilliantly! #AlphaGo thought it
was doing well, but got confused on move 87. We
are in trouble now...
Mistake was on move 79, but #AlphaGo only came to
that realisation on around move 87
When I say 'thought' and 'realisation' I just mean the
output of #AlphaGo value net. It was around 70% at
move 79 and then dived on move 87
Lee Sedol wins game 4!!! Congratulations! He was
too good for us today and pressured #AlphaGo into
a mistake that it couldn’t recover from
From: https://twitter.com/demishassabisGuys, please, publish charts of win prob estimated by alpha go in time during these games. Some heatmap telling which moves did it consider as best for both sides during the games would also be cool, but that's surely more time consuming to prepare.
It would be great to be able to have such things for top pro tournaments in the future.
[1] http://www.nature.com/nature/journal/v529/n7587/fig_tab/natu...
I don't require the screenshot to be instantaneous, I require it to appear instantaneous. In that sense, if the whole rendering pipeline is working to give me a framerate of 60 fps, then I could spend ten times as much rendering a screenshot without that delay being noticeable. Also, why on earth does Windows (don't know the behaviour on Linux/Mac) apply ClearType to a screenshot? That has always bugged me, there are some situations where you tolerate it and others where it hurts.
I felt so sorry for Lee Sedol when I saw him lose the second match, facing an empty chair ,and he could only ask one of his friend to review the game.
Generally I'd guess you're right though.
I'm not sure how you define "deranged" (AFAIK, that's not a medically defined term within the DSM-IV) but most of those brilliant people end up being a little 'off'. My father was an academic, one of the people he went to graduate school with was working in Boston while I was a child. He was absolutely groundbreaking work but he's so difficult to collaborate with (think: the mannerisms of Richard Stallman) that he's been floating around universities until his welcome is worn out. He can figure out remarkable things in higher level computational chemistry, but he can't really figure out humans. Had he decided instead during the 80s to work at Renaissance instead of pursuing academic research, he almost certainly would be worth in the low hundreds of millions.
For example, my brilliant programmer friend that can't hold a job. Or the artist that can only paint their own inspirations. Or the savant that doesn't get along with anybody.
A Chess or Go champion is probably doing the single-most lucrative activity that they are capable of.
What the person you're responding to means is that if they have the mental capacity (and also discipline) to play the game at such a high level, they could probably also excel in other professions if their goal was to make money instead of doing something they loved.
Of course we can debate whether they would achieve such success if they were doing something they didn't enjoy as much, but I think most people would argue that most chess or go champions could make more money or in other words do more "lucrative" activities if they chose to do them instead of dedicated all their time to the game.
You seem to be arguing against the idea, "A chess grandmaster is a genius and therefore can walk into Google and immediately start doing more and better work than most of their senior programmers". I don't think anyone believes this (correct me if I'm wrong).
I think a more serious idea is, "A chess grandmaster is a genius and therefore learns faster and has a higher performance ceiling and such than most people, and if they spent a couple of years learning to program, they could become an entry-level Google programmer, after which they would rise more quickly than most hires, and eventually would outperform most of Google's senior programmers."
By the way, I think most of the best chess players were extreme chess prodigies. (Just looked at Kasparov, Karpov, Shirov, Kramnik, Anand, and Carlsen's Wiki pages; all but Kramnik had the year listed, and they all became grandmasters around age 17-19, except Carlsen, who was around 14. Kramnik's page mentioned winning a gold medal for the Russian team at age 16, and that he wasn't a grandmaster when selected for the team and this was unusual.) I think this is consistent with them being highly gifted children, who choose to spend their time doing chess.
While not exactly what you wrote, people do think that someone is a genius at some subject can become a genius at another area with less work than it took for someone in either field to originally become a genius at that field, especially society groups the areas together (so sports star becoming master programmer is far less likely to be believed than chess master becoming master programmer).
Do you have a reference for that please? I'm interested in the subject.
Of course chess grandmasters aren't all-around geniuses, but many traits that are prerequisites for a successful chess player certainly are translatable into careers in other fields, and correlate with above average brainpower (good memory, discipline, long attention span, spatial intelligence etc.)
Botvinnik was an accomplished engineer, Euwe had PhD in mathematics, Anatoly Karpov is a millionaire, interestingly enough a lot of recognized chess players had careers in music (like Taimanov or Smyslov)...
Being a grandmaster surely requires an above average intellect, there is no such thing as a chess savant. While it's an urban legend that Kasparov's (arguably the greatest chess player ever) IQ was in the ballpark of 190, he did clock at 135. Such a result is not unheard of, yet still placing him in top 1% or so.
> In a recent review, Ericsson and Lehmann (1996) found that (1) measures of general basic capacities do not predict success in a domain, (2) the superior performance of experts is often very domain specific and transfer outside their narrow area of expertise is surprisingly limited and (3) systematic differences between experts and less proficient individuals nearly always reflect attributes acquired by the experts during their lengthy training.
I will say this for Chess or Go grandmasters: They have drive and dedication (and in some cases, compulsiveness). That alone would probably allow them to do better than the average person in another field if they had pursued that field from the get-go. Also, I'd caution you against relying on hand-picked anecdotes; I could just as easily pick out a bunch of Chess players who weren't good for anything else. You'd need broad-based statistics.
1. A champion Go/Chess player could retire, and jump a successful career as <Something else>. You refute 1.
2. A champion Go/Chess player could have pursued a career in <Something else> from the start and made more money. You have not refuted 2.
For an example, many physics PhDs become successful software engineers and quants, making a mid-career jump to a different field.
Not really refuting your statement, which is essentially true, but it's worth your time to read the wikipedia page.
He played on a professional hockey team with his two adult sons at one point(!). He played in the NHL in five different decades.
Julio Franco played baseball till 49.
George Blanda played football till 48
Nat Hickey played baseball till 45.NFL seasons are 16 regular season games, plus 4 pre-season games. And then there are the playoffs, which a given team might or might not make or advance in.
All told, given that the signing bonus is amortized over all the games he plays, the annual salary, etc., I think it would be fair to say that Vernon will make around a million dollars a game.
Aside: Vernon isn't necessarily "the best" DE in the NFL, but due to market forces and the way things work with the salary cap, free agency rules, etc., the contract he just signed is one of the largest for a defensive player in the league. QB's tend to make even more, but I can't recall a really high profile QB who has signed a big deal recently.
[1]: http://www.spotrac.com/nfl/new-york-giants/olivier-vernon/
A really big issue with NFL contracts is this "guaranteed money" thing...I believe the NFL is the only major US sports league who give player contracts without it, so you have to take those salary numbers with a grain of salt.
I would like to see AlphaGo play 21 against Steph Curry.
When I see "-Unknown", I interpret it as "somebody said it, but no one is quite sure who".
"(NN)" in the current thread was used to indicate "I personally don't know", which does not imply the quote is unattributable.
NN can be used in general for people whose origin is uncertain. I think in this specific case it's a bit misleading - although the wikipedia article seems to suggest that NN can also be used as a synonym for "unknown", although from a historical perspective it is a bit incorrect.
The literal translation from Latin creates some confusion if you didn't know the context.
I hope this helps.
I hope your reductionist explanation is not accurate because it would imply that we are already so highly conditioned to machines and to AI that this match is thought to be no different from a match between two humans.
Imagine what AI, in different fields, means for humanity if it has so much to teach us, just by being able to "think" differently. I sure hope that one day they start writing philosophy and, doing so, potentially legitimize everything that currently makes humans unique and extraordinary.
(also during the actual game with more details, but I can not find the exact time again). Edit : can not find game3 commentary but during game 4 he is coming back to it again with interesting details as for why this is a great opportunity to inspire human players : https://youtu.be/yCALyQRN3hw?t=7097
Basically, once in ancient Japan and more recently a Chinese origin player in Japan, both incredibly strong players, surprised everyone with never before seen moves that subsequently where integrated in modern game theory. The hope is that the same could happen here.
Dolphins are about as intelligent as us, too. Are dolphins amoral? Do they delegitimize Beethoven, Tesla, Gödel, Einstein, and da Vinci?
http://www.deepseanews.com/2013/02/10-reasons-why-dolphins-a...
While it's somewhat news to me, I'm very glad to hear we've moved past that.
I have no reason to doubt that quantitatively we have made some progress - although Pinker's arguments are most definitely not universally accepted in the scientific community.
/s
If we invented dolphins, and they began to replace and displace us, then maybe? I don't think it is about the intelligence, but in how we use it of course. Dolphins don't really have any control of my life or those around me.
That sounded a lot more paranoid than I meant, I actually agree with what I think you are saying
How are you making this comparison?
That's quite an accomplishment.
And for a game that was thought to be many decades away in terms of computer capabilities, this probably should be a time to think of the possible consequences of AI improvement and take a close look at it.
That line of reasoning would be worth a laugh, if it wasn't so widespread in the general population.
>Both have to do with right and wrong, but amoral means having no sense of either
Well that's true then! A nice quote is from Feynman in regards to scientific research.
"To every man is given the key to the gates of heaven; the same key opens the gates of hell. — Richard Feynman "
We really have no idea yet how this AI research will play out.
I don't quite like the line "Delegitimize all that makes people unique and special" But that's my only real disagreement.
If you care about being unique and extraordinary more than about reason, knowledge, truth, the observable reality, and the search for what it really means to be sentient, then and only then you may call AI "amoral".
Also, you are being racist against artificial sentient beings, and being racist is hopefully not what makes humans extraordinary.
Q: We heard there's now an anti-Lee Sedol website in Korea?
A: I don't even have time for my fans. I don't care about haters. ("나를 좋아하는 팬들에게도 신경을 못 쓰는데 그들에겐 당연히 신경 끈다.")
Or, writing from the point of view of our mechanical successors to this world, for helping to advance a highly ethical field that could exterminate that genocidaly murderous evolutionary abomination that was the human race. Who incidentally thought they were extraordinary but couldn't even play Go.
You're treating Lee Seedol as if he is a fixed dataset to be trained on. Why can't he also be a "NN" who can also "refine" his AI[0] and thus be hearder for Alpha Go to compete with?
You're putting too much faith in the machine and dropping the person, that might not be completely fair.
[0] Albeit not artificial in this case.
That's not factoring in other information, like Sedol now being familiar with alphaGo's strategies and improving his own strategies against it.
So there is a good chance he is now evenly matched with AlphaGo, and likely much better than the single machine version.
On the other hand, the Nature paper shows the single 8 GPU machine performs similar to the 64 GPU cluster, but the larger clusters perform a comfortable margin better. [0]
By a single machine winning many games relative to the distributed version, really it's just saying that the value/policy network is more important than the monte carlo tree search. The main difference is the number of tree search evaluations you can do; it doesn't seem like they have a more sophisticated model in the parallel version. The figure suggests that there are systematic mistakes that the single 8 GPU machine makes compared to the distributed 280 GPU machine, but MCTS can smooth some of the individual mistakes over a bit.
[0] http://www.milesbrundage.com/uploads/2/1/6/8/21681226/877172...
A human is far, far, leagues, more efficient at learning than today's AIs. These AI requires millions of hours of man time of data to even come close to competing at the level of an expert person which did the same, and even arguably far better, in a "few decades".
It took some serious hardware for Deep Blue to defeat Garry Kasparov. Now there are smartphone apps with that same level of Chess-playing skill. And if anything the AlphaGo approach is more amenable to running on lower-specced hardware without requiring help from Moore's Law (because you simply train it better).
Hassabis and Silver kind of reminded me of developers given the details of a bug that was notoriously difficult to find.
edit; I can't wait for the reviews of the entire 5 game series. If a book came out I'd very likely buy it. A book discussing both Go and AlphaGo AI at the same time consisting of people among the top of their respective fields would be amazing.
> q
> quotes like this
since it's markdown - often you get commentish things that allow limited HTML plus indented code blocks, and people get used to using the latter as the only thing that works everywhere
HN does not use Markdown.
Pre is broken on mobile, forcing users to scroll a horizontal line which is incredibly annoying if it's a long line.
Here's a screen shot: http://imgur.com/Ru56wMKAnd pre with very long unbroken lines is even worse.
I agree with you that it's annoying when people post long indented lines.
And no, HN does not use Markdown.
It isn't a huge time sink to wrap a paragraph in stars.
I agree with you, it's annoying.
I never like to assume malice, but ...
The bad moves in the eyes of humans could be risky bets or the horizon effect. [1] [2] AlphaGo's use of deep neural nets (value networks) to evaluate board positions should significantly help counter the horizon effect, but since move 78 by Lee Sedol which turned the situation around was unexpected by some top pros (Gu Li referred to it as the 'hand of god' [1]), the patterns which follow are likely rare in possible game states and therefore not strongly embedded into the value networks, leading to AlphaGo's loss.
I hope the DeepMind team will help enlighten us on this in the near future.
[1] https://gogameguru.com/lee-sedol-defeats-alphago-masterful-c...
If anything, it seems to me that AlphaGo's problem here might be the time management: Seeing a really scary situation, Lee Seido just sank many minutes into reading the problem, going pretty much all the way to byoyomi time. A human, after seeing something like that, would figure out that their assessment of the situation and their opponent's is very different, and spend a lot of budget trying to figure out what was wrong. AlphaGo just didn't see the problem, and didn't just budgets its time to analyze the position to death. It moved slower than before, but not really that much, and ended up making moves a kyu player could see as terrible.
Either way, I'd love to see Deepmind giving us all a good postmortem of the 70-100 range of moves.
Beyond that, I think that AlphaGo may still be missing a type of component. From the descriptions of it, the policy network generates possible moves from board positions, and the value network evaluates the probability of desirable outcomes. How this is different than human play is that strategic assessment and planning are implicit in the middle layers rather than a 'conscious' element to searching and decision making. I'm not saying that this is a necessary component as AlphaGo has already done exceptionally well. I do believe this kind of 'middle-out' processing producing and evaluating strategic concepts could make it better handle unusual circumstances. Being trained on high amateur and pro games, it will best respond to the most conventional of those types of games, more unconventional the game becomes, the worse it would fare in terms of efficiency of move generation and choices of which to evaluate.
I suppose beating humans isn't AlphaGo's primary motive though - learning to play a perfect game of Go in general is probably more difficult than playing the perfect game against a particular person.
The AI player can't give up information in this manner because it lacks eyes, so I'd say that it should not be able to use this information from the human player.
Of course, a plausible alternate explanation is that AlphaGo felt like it needed to make risky moves to catch up.
When you have a game of Go, or Super Mario level. You don't want to make your decisions by just checking the local features and doing them, because it can be the case that by compounding errors you end up in a state you never saw, and all of the future decisions won't be good.
One can avoid these situations by training jointly over the whole game.
For example, maximum entropy models can work for decision making problems but their training leaves them in a "label bias" state because the training is trying to minimize loss of local decisions, instead of trying to minimize the future regret of current local decision.
The solution to these label bias problems are Conditional Random Fields, or Hidden Markov Models. You could accomplish the same with Recursive Neural Networks if you trained them properly. For example, there is no search part (monte carlo tree search, or dynamic programming [viterbi] like it is in CRFs or HMMs) in RNNs but they are completely adequate for decision based problems (sequence labeling etc.). Why is that the case? Because search results are present in the data, there's no need to search if you can just learn to search from the data.
If DeepMind open-sourced the hundreds of millions of games that AlphaGo played, it is quite possible to train a model that wouldn't need a Monte Carlo search and would work quite well, because you would learn the model to make local decisions to minimize future regret, not to minimize its local loss. [1]
The only reason why reinforcement learning is used is because there are too few human games of Go available for the model to generalize well. Reinforcement learning can be used in the setting of joint learning because you play out the whole game before you do the learning. This means that you can try to learn a classifier that will minimize the regret by making a proper local decision. Although, as far as I know, and can see from the paper, they didn't train AlphaGo jointly over the game sequence.
But! Now they have a lot of data and they can repeat the process.
[1]: http://arxiv.org/abs/1502.02206
[2]: http://repository.upenn.edu/cgi/viewcontent.cgi?article=1162...
I fail to see how a "per-state normalization of transition scores" translates to there being a bias in value networks towards states with fewer outgoing transitions.
Their value policy network isn't trained jointly and can compound errors. There are approaches with deep neural networks that don't have a joint training but work pretty well. The reason is that networks have a pretty good memory/representation and by that they avoid much of the problems. But for huge games like Go it is quite possible that more games need to be played for these non-structured models to work well.
The concept of label bias, or decision bias is a joint/structured learning concept. It is a machine learning concept, it has nothing to do with the application. There are training modes with mathematical guarantee that the local decisions will minimize the future regret.
Joint learning is done not on the whole permutation but on the markov-chain of decisions, which is sometimes a good enough assumption. For example, the value policy network of AlphaGo is percisely a Markov chain, given a state, tell me which next state has the highest probability of victory. The search then tries to find the sequence of moves that will maximize the probability, and then it makes the best local decision (one move). It works like limited depth min-max or beam search. They do rollouts (play the whole game) to train the value network, but it is now a question if they train it to minimize the local loss of the made decisions, or if they train it to minimize the future regret of a local decision. As I've stated before, minimizing joint loss over the sequence, or minimizing local loss over each of made decisions, is exactly influencing if there will be bias or not.
The whole point of reinforcement learning is to create a huge enough dataset to overcome the trajectories-not-seen problem. The training of the models for playing Go is entirely a whole different kind of a problem.
Now when they have hundreds of millions of meaningful games they can skip the reinforcement learning and just learn from the games.
The illustration of the "label bias" problem is available in one source I referenced. Terms like compounding errors and unseen state are there. The "label bias" is present only in discriminative models not generative ones. Which means that AlphaGo - being a discriminative model, can suffer from "label bias" if it wasn't trained to avoid it.
The compounding errors problem that stems from decision bias isn't because you haven't seen the trajectory, it is because the model isn't trained jointly.
We're talking about the same thing. You just aren't familiar with the difference present between joint learning discriminative models and local decision classifiers (Markov entropy model vs conditional random fields - or recursive CNNs trained on joint loss over the sequence or recursive CNNs trained to minimize the loss of all local decisions).
In the case of Go, one would try to minimize the loss over the whole game of Go, or over the local decisions made during the game of Go. The latter will result in decision bias - that will lead to compounding errors. The joint learning has a guarantee that the compounding error has a globally sound bound. (proofs are information theory based and put mathematical guarantees on discriminative models applied to sequence labelling (or sequence decision making))
edit:
Checkout the lecture below, around the 16 minute mark it has a Super Mario example and describes exactly the problem you mentioned. The presenter is one of leading figures in joint learning.
https://www.cs.umd.edu/media/2015/12/video/17235-daume-stuff...
It is completely supervised learning problem. But, look at reinforcement learning as a process that has to have a step of generating a meaningful game from which a model can learn. After you have generated bazillion of meaningful games you can discard the reinforcement and just learn. You now try to get as close to the global "optimal" policy as you can, instead of trying to go from an idiot player to a master.
Of course, the data will have flaws if your intermediate model plays with a decision bias. So, instead of training the intermediate to have a bias, train it without :D
Although, if you checkout his papers, the problems I've talked about, when you have more than enough data and when you know you should be able to generalize well you still can get subpar performance if you don't optimize jointly. AlphaGo model isn't optimizied jointly but its power mostly lies in the extreme representation ability of deep neural networks.
They maybe could have anyway, but it would be cheating: just the same as if they'd Mechanical Turk'ed it by e.g. having Ke Jie actually choose the moves to play.
Of course against a really strong player you're going to get beaten after that but a weak player strong on theory will have a harder time.
And for a complex system transitions happens to be sensitive to the conditions and with quite a lot of impredictability.
AI cannot do smooth transitions because they lack the intuition of what smooth means, and that's how to win against them.
1) identify a basin of attraction(apparition of a bounded domain of evolution in a space phase) 2) set the AI in a well known basin by tricking it; 3) imbalance the AI by throwing garbage behaviour that kick him out of the basin in a random direction 4) let the human win in the chaos that ensues.
Of course it is better done with a software to help you. A real time spase phase analysor.
The point is like in a lot of domain, construction of an AI requires more energy than a software for sabotaging it.
But once you get the framework of thoughts to win against an AI you can get all the AI.
Demis Hassabis said of Lee Sedol: "Incredible fighting spirit after 3 defeats"
I can definitely relate to what Lee Sedol might be feeling. Very happy for both sides. The fact that people designed the algorithms to beat top pros and the human strength displayed by Lee Sedol.
Congrats to all!
So it's not totally arrogant of Ke Jie to suggest he could beat AlphaGo. AlphaGo has not much 'experience' in dealing with Ke Jie.
And Lee winning of game 4 shows a human is still indeed more capable than any AI. He basically reprogrammed his game on his own. Sorta.
Came up because of an assertion from the interviewer that chess AI had been trained against it's opponent specifically.
Personally I want to see a discussion game with the top Go experts (including both Lee Sedol and Ke Jie amongst others) competing against the next version of AlphaGo, in a game with much longer time limits.
So Lee's games are sort of a drop in a bucket as far the performance of the AI goes.
I think it's premature, establishing bounds with good confidence interval requires tens or hundreds of games. Specifically, 3:2 result would be really inconclusive.
To use an analogy, having confirmation of contact by even a single alien species would be hugely important, way more so than exactly nailing down the number of alien species. Knowing that something is even possible is oftentimes the most important aspect that needs to be ascertained, and contact (or a win, in Lee's case) does that unequivocally.
If Chess were truly solved, then you wouldn't be able to make a new AI program that could do better than even odds against the existing ones. But that's not the case, and incremental advancements in Chess-playing programs are made all the time. There are even tournaments where Chess programs play each other. If Chess were solved, such a thing wouldn't make any sense, just like how there are no Tic-Tac-Toe tournaments because that game is solved.
If chess were _solved_ we'd know a strategy to allow one of the players (likely white) to always win, or for either of them to always force a draw. (and we'd know which of these strategies were possible for chess).
Consider, say there is a first move white could choose such that no matter what moves black makes, white will win. Then the first would be true. Instead consider, that there is no such move, and any first move has choices where either could win-- if some of those are ones which would force a black win, then the first is again true but for black. Otherwise, a draw can always be forced. These are the only possible outcomes for a solved game of chess.
Of course, what's obvious to a human might not be so at all to a computer. And this is the interesting point that I hope the DeepMind researchers would shed some light on for all of us after they dig out what was going on inside AlphaGo at the time. And we'd also love to learn why did AlphaGo seem to go off the rails after this initial stumble and made a string of indecipherable moves thereafter.
Congrats to Lee and the DeepMind team! It was an exciting and I hope informative match to both sides.
As a final note: I started following the match thinking I am watching a competition of intelligence (loosely defined) between man and machine. What I ended up witnessing was incredible human drama, of Lee bearing incredible pressure, being hit hard repeatedly while the world is watching, sinking to the lowest of the lows, and soaring back up winning one game for the human race.. Just incredible up and down in a course of a week. Many of my friends were crying as the computer resigned.
Just letting others know, this expression is a rather common way of saying a single move that changed the course of the game. There were "divine-inspired moves" that AlphaGo made in the first three games too.
Add to that the moves where AlphaGo basically threw away stones by adding to formations that would be removed from the table. Even I, a complete, lousy, amateur, could see that they were a mistake.
Training an ai to make good play in a bad situation would require it to train in ways that are very different than the AlphaGo vs AlphaGo training that it spent a lot of time doing. And why do that, instead of trying to make itself good while the game is even, or when it's winning?
It's a bit like how it's different to train in chess to play in pro games, vs training to hustle amateurs in the park: You are not making the best move, but a good move that will confuse the opponent the most. You are trying to exploit a bad opponent: Very different play.
Toward the end AlphaGo was making moves that even I (as a double-digit kyu player) could recognize as really bad. However, one of the commentators made the observation that each time it did, the moves forced a highly-predictable move by Lee Sedol in response. From the point of view of a Go player, they were non-sensical because they only removed points from the board and didn't advance AlphaGo's position at all. From the point of view of a programmer, on the other hand, considering that predicting how your opponent will move has got to be one of the most challenging aspects of a Go algorithm, making a move that easily narrows and deepens the search tree makes complete sense.
If I had to guess (and this is pure speculation), AlphaGo has no concept of waiting for its opponent to make a mistake. Instead, it assumes its opponent will continue to make the best possible follow-ups, and so AlphaGo feels overly compelled to "keep up". In this case, that did it in.
If this is what happened, then yes, I would expect Lee to be able to capitalize.
This has also resulted in larger shifts in playing style over time. Studying very old (and I mean very old...700+ years old) games can be entertaining and even educational in the abstract, but you won't want to directly adopt the style of play because the game has evolved.
It's already been mentioned a couple of times that AlphaGo almost certainly represents just such a shift. Top players will learn from it, and I'd even be willing to bet they will beat it with some regularity once they do!
Ultimately, what sets apart Go geniuses is their ability to play creatively in the face of seemingly insurmountable challenges. So the big question is how "creative" AlphaGo can be. Is it merely synthesizing strong play from known positions? Can it introduce novel strategies? And if it does, will it be able to adjust as other Go masters adjust to it and bring their own brand of creativity to play?
To answer your original question, this very well could introduce a new era of more aggressive play to the world of Go. Only time will tell...
This incarnation of AI is not creative, it wont generate new play styles, that is still the domain of top human players for now. But it will ruthlessly learn and adopt any new and improved strategies. That's really the point to take away from its success so far.
The OMG-AI people claim that AGI would be dangerous because it would reliably innovate in new spaces and out-predict humans.
So a true super-AGI would make go moves that were unexpected and incomprehensible with some percentage of misleading fake-outs, but it would still win most or all of the time.
If the human exploration of Go-Space is close to the god's hand bounds, this can't be true.
We'll know if this is the case in a couple of years, if the competition between human and AI goes back-and-forth (unlike Chess, where after AI was good enough to beat humans, it could do so reliably).
Either way, it's interesting to note that AlphaGo had literally thousands of games to learn from to find weaknesses in human play, but Lee Sedol seems to have only needed 3 before he was able to find weaknesses in AlphaGo's play.
To be fair we can't know how many games Sodol played in his own head to figure this out.
But they're already working on a new version of AlphaGo which isn't trained on any human data at all. It starts by making truly random moves and improves from there. This will require much more processing time and probably an order of magnitude more "self-play", but it will probably result in truly novel strategies that aren't part of the current human metagame.
One of the DeepMind guys just confirmed that this is how AlphaGo operates in the press conference.
-- Pyanfar Chanur (C.J. Cherryh)
This is really interesting, because forming a model of our opponent and tailoring our strategies appropriately is fundamental to how humans approach competitions.
In other words, the humans commentating the game were evaluating the moves as non-sensical because the outcome (AlphaGo plays here so Lee plays here) is a foregone conclusion and doesn't change the human evaluation of the board position. What it does do is remove uncertainty (AlphaGo plays here, but Lee screws up and plays somewhere else). In their evaluation, humans tend to value that uncertainty (i.e. counting on the possibility of a mistake), but I'd guess that AlphaGo penalizes the uncertainty (i.e. known board positions are scored higher than potential board positions), leading it to over-value simple advancement of the board in the end-game.
Maybe its training was focused on winning, not loosing narrowly. So as soon as it became obvious it can't win. It was just making silly moves because it was less researched scenario.
Some people when they see they can't win they do silly moves just for fun.
The moves made by AlphaGo there were very bad.
Go programmers have taken various steps to mitigate this behavior, such as dynamically adjusting komi to trick the engine into thinking it is a closer game, but I don't know if AlphaGo uses any such technique.
Another interesting thing I noticed while catching endgame is that AlphaGo actually used up almost all of its time. In professional Go, once each player uses their original (2 hour?) time block, they have 1 minute left for each move. Lee Sedol had gone into "overtime" in some of the earlier games, and here as well, but previously AlphaGo still had time left from its original 2 hours. In this game, it came down quite close to using overtime before resigning, which is does when the calculated win percentage falls below a certain percentage.
Tesuji isn't a trick play, it's more like a power play. Each player can read out how a fight is going and see their line far into the future. Two professionals will pick two lines, two suji, which are in balance and push up against one another tightly.
A tesuji is a part of the line which is suddenly showy or strong. It could mean a failure for the opponent if they had not taken enough of an advantage in the struggle to this point or if they do not have a counter tesuji available.
Indeed, that might be the design of a set line: one side continually loses ground to the other forcing the other to take these small advantages all so that the first side has an opportunity to play a tesuji and return to balance. Many such lines are canonicalized ("joseki") and known to any professional. Moreover, professionals regularly identify potential tesuji and expect their opponents to as well.
On one hand, we have racks of servers (1920 CPUs and 280 GPUs) [1] using megawatts (gigawatts?) of power, and on the other hand we have a person eating food and using about 100W of power (when physically at rest), of which about 20W is used by the brain.
[1] http://www.economist.com/news/science-and-technology/2169454...
Probably on the order of one megawatt or so.
http://inhabitat.com/infographic-how-much-energy-does-google...
1920 CPUs (a 4-core haswell from 2013 is around 170Gflops). 280 GPUs (previous gen Nvidia K series peaks at around 5200GFLOPS). That's 1,782,400Gflops or around 150,000x more processing power. If they were running latest-gen hardware, then the would be closer to 200,000x faster.
Given that Moore's law is slowing down and the size of the system, we're a long way from considering putting that in a smartphone.
AlphaGo is still a very new program (two years since inception). It will get significantly better with more training, or, equivalently, it will stay at the same strength while running on much less hardware.
Don't read too much into what one particular snapshot in its development cycle looks like. Humanity has had hundreds of millions of years to maximize the efficiency of the brain. AlphaGo has had two years. It's not a fair comparison, and more importantly, it's not instructive as to what the future potential of AI algorithms looks like.
The result "W:Resign" was added to the game information.
Edit: Tinyyy is right.
[1] http://gall.dcinside.com/board/view/?id=baduk&no=109200&page...
https://github.com/lukaszlew/EasyGoGui/blob/master/src/net/s...
And search for MSG_RESIGN_2
off-topic: DeepMind should switch to a tiling window manager like i3 for increased keyboard-only productivity :)
But there is nothing wrong with keeping the frontend machine used in this Go match in default Ubuntu desktop environment since its only purpose is to play Go with a graphical user interface anyway.
Also given the progress of DeepMind so far, it's very likely that whatever desktop setups they have, is working very well for them.
The contra argument for the sake of the argument goes, he just was lucky to find a local maximum (?) outside of the search space (?) by chance, rather than learning in a few days the universal function (?) that the NN thinks solves go, or at least one fixed point (?) [i.e. the surprisingly wrong expectation].
I tried to estimate it mathematically. Using a uniform distribution across possible win rates, then updating the probability of different win rates with bayes rule. You can do that with Laplace's law of succession. I got a 20% that Sedol would win this game.
However a uniform prior doesn't seem right. Eliezer Yudkowsky often says that AI is likely to be much better than humans, or much worse than humans. The probability of it falling into the exact same skill level is pretty implausible. And that argument seems right, but I wasn't sure how to model that formally. But it seemed right, and so 90% "felt" right. Clearly I was overconfident.
So for the next game, with we use Laplace's law again, we get 33% chance that Sedol will win. That's not factoring in other information, like Sedol now being familiar with AlphaGo's strategies and improving his own strategies against it. So there is some chance he is now evenly matched with AlphaGo!
I look forward to many future AI-human games. Hopefully humans will be able to learn from them, and perhaps even learn their weaknesses and how to exploit them.
Depending how deterministic they are, you could perhaps even play the same sequence of moves and win again. That would really embarrass the Google team. I hear they froze AlphaGo's weights to prevent it from developing new bugs after testing.
Also, he won with white but he will play with black next time, so playing the same sequence of moves can't happen. Additionally, even if the AI didn't incorporate any randomness in the opening (I think it does) it may choose different moves if it gets a different amount of time to think, so Lee Sedol would have to play his moves at exactly the same time as the last game. A couple of seconds deviation only has to lead to a different move in one of the 80 or so moves before the mistake was made to invalidate this strategy.
Most importantly, they chose this point in time because they know that there are several other AI research teams that are also on the right track (including at Facebook). Like circumnavigating the globe or landing on the Moon, 90% of the benefit of it is lost if someone else does it before you. So you don't wait for 100% certainty of winning -- you are maximizing for the chance of winning first, not winning for certain.
Given that, it seems reasonable that they would go after Lee Sedol when they were sure they were better than him, but not too much better than him. So a non-5-0 outcome is, in hindsight, not horribly surprising.
Alphago was an entirely new method, using deep convolutional neural networks as a move generator. Therefore there wasn't any guarantee that it had to be just slightly better than previous Go playing algorithms, it could have easily been far above humans.
Likewise this is also the first time Go AI has been given Google scale resources. Both in terms of a team of the best researchers working full time on it, and their computing power. Whereas previous Go projects were all hobbyist things.
And Google didn't wait 10 years to challenge Sedol. The match was arranged only a few months at most after they started developing it.
Didn't they say that it's not considered a "bug" but rather how AlphaGo "thinks"? "when it's winning it doesn't care about how much it's winning, and when it's losing it doesn't care how bad it's losing"
When you're winning, a good move has a mathematical definition; it's a move that, given optimal play by both sides, will lead to victory for you. Computers aren't powerful enough to be able to calculate exactly what moves those are, but it's at least well-defined in a way that they know what they're looking for.
When you're losing, there's zero moves that, given optimal play by both sides, will lead to victory for you (so all moves are equally "bad" in a mathematical sense). Instead, "attempting a comeback" involves hoping your opponent will mess up some way, so in that sense, what constitutes a good move isn't mathematical but more about predicting how your opponent thinks and where they might mess up.
AlphaGo has mostly trained by playing itself, so the ways it thinks its opponent might mess up are probably completely different from how an actual human messes up.
Another possibility is that it was looking way deeper than anybody, and there was a 1% chance or turning the whole game around with those seemingly bad moves. But Lee Sedol blocked that deep move in a way nobody was able to see.
Myungwan Kim in his commentary, estimated the game to be worse for white even after 79 if Black had not destroyed it's opportunity for fighting the ko at M-13 - https://www.youtube.com/watch?v=SMqjGNqfU6I&t=1h40m1s
Edit: here's another great one on MCTS: https://gogameguru.com/alphago-4/#comment-13479
Is there something I'm missing?
For all we know AlphaGo has perfectly fit amateur games, but professional games are on a whole different level
You tell whether overfitting is a problem by evaluating performance on a held-out test set.
From the cited paper,
> [Experimental results] suggest that adversarial examples are somewhat universal and not just the results of overfitting to a particular model or to the specific selection of the training set.
Anyway, Monte Carlo Tree Search is bad at losing positions. In general, you want to delay the impending catastrophe as long as possible instead of making stupid moves that make your position worse and worse. However, MCTS uses random rollouts to the end of the game, which sometimes make it difficult to ascertain if the inevitable doom is near or far.
Also, MCTS converges very, very slowly and is likely to miss a unique, single winning continuation.
I think it is probably a combination of both AlphaGo's value network failing to realize the good position of Lee Sedol after his brilliant play, and the MCTS failing to spot the unique winning sequence for Lee, that caused it to make the mistake. But we should probably wait for official analysis from the Deepmind team to see what exactly went wrong.
[0] "Intriguing properties of neural networks" http://arxiv.org/pdf/1312.6199.pdf
One thing I've been pondering is if many adversarial samples exist. The board is rather low dimensional (19 x 19) and discrete. While certainly a massive state space, one of the suggestions for why adversarial images work is that the real number line is incredibly dense.
For example our possible Go board space is 2^(log2(3) * 19^2) for Go but 2^(24 * 28^2) for greyscale [0, 1] normalized single precision float imagery for MNIST. Thats an exponentially bigger space (I think something like 1e910 times bigger!), and gets only larger if you train with double precision, have larger images, add multiple color channels, add more nonlinear layers, etc.
Lee Sedol won because he played extremely well. But when AlphaGo was already losing it made some very bad moves. One of them was so bad that it's the kind of mistake you would only expect from someone who's starting to learn how to play Go.
On the surface, as an analogy, it sounds like investors in financial markets, capitulating, selling at a loss for a more risky outcome. In hindsight almost always bad moves, but at the time of making them it feels right because it's removing risk. Investors are losing, and then when capitulating they make even worse moves, like selling at market bottoms.
http://www.investopedia.com/terms/c/capitulation.asp?layout=...
I think the answer would be most likely not - the monte carlo tree search is randomized so AlphaGo's responses to Sedol may not be exactly the same, requiring Sedol to not be able to repeat the exact same play.
Now if that same picture held for Go, then a situation like this would seem to be impossible. Either the computer should be much worse than a human player, or much better. It would be an incredible coincidence that, at the end of six months of training, the computer happened to be of comparable skill to humans.
For the game of Go, at least, Yudkowski is wrong. What other aspects of intelligence are this way? Yudkowski's picture seems appealing, but perhaps it is wrong for many areas of intelligence.
In the linked article, Yudkowsky even says "On the right side of the scale, you would find Deep Thought—Douglas Adams's original version, thank you, not the chessplayer." The implication is clear that these programs playing chess/Go are nothing like what he is talking about - general AI.
Or so I assume, from my less-than-complete understanding of Yudkowsky's writings.
I don't think it's that incredible - By 18 years old a significant proportion of high school students know more about chemistry than the best scientists up to 1800 did, combined.
There's a lot of human games for AlphaGo to look at, but if it is to exceed human level of play, it'll have to figure how to do that by itself. Look at human level games, learn to play human level.
It's quick to get to the edge of human knowledge, and slower to go beyond it.
I don't believe AlphaGo had the time to do any additional training between matches. So effectively Lee has the ability to 'learn his opponent' while AlphaGo cannot until the entire match set is over because of how long it would take do do additional training.
A living brain is made of neural nano-processors.
The author thinks that Lee Sedol was able "to force an all or nothing battle where AlphaGo’s accurate negotiating skills were largely irrelevant."
[...]
"Once White 78 was on the board, Black’s territory at the top collapsed in value."
[...]
"This was when things got weird. From 87 to 101 AlphaGo made a series of very bad moves."
"We’ve talked about AlphaGo’s ‘bad’ moves in the discussion of previous games, but this was not the same."
"In previous games, AlphaGo played ‘bad’ (slack) moves when it was already ahead. Human observers criticized these moves because there seemed to be no reason to play slackly, but AlphaGo had already calculated that these moves would lead to a safe win."
Which, I add, is something that human players also do: simplify the game and get home quickly with a win. We usually don't give up as much as AlphaGo (pride?), still it's not different.
"The bad moves AlphaGo played in game four were not at all like that. They were simply bad, and they ruined AlphaGo’s chances of recovering."
"They’re the kind of moves played by someone who forgets that their opponent also gets to respond with a move. Moves that trample over possibilities and damage one’s own position — achieving less than nothing."
And those moves unfortunately resemble what beginners play when they stubbornly cling to the hope of winning, because they don't realize the game is lost or because they didn't play enough games yet not to expect the opponent to make impossible mistakes. At pro level those mistakes are more than impossible.
Somebody asked an interesting question during the press conference about the effect of those kind of mistakes in the real world. You can hear it at https://youtu.be/yCALyQRN3hw?t=5h56m15s It's a couple of minutes because of the translation overhead.
I wonder if Lee Sedol can find a way to replicate that in Game 5.
[0]: https://twitter.com/demishassabis/status/708928006400581632
https://www.youtube.com/watch?v=yCALyQRN3hw
At the end, Lee asked to play white in the last match, and the Deepmind guys agreed. He feels that AlphaGo is stronger as white, so he views it as more worthwhile to play as black and beat AlphaGo.
Conference over, see you all tomorrow.
https://www.youtube.com/watch?v=yCALyQRN3hw&feature=youtu.be
AlphaGo made a mistake and realized it was behind, and crumbled because all moves are "mistakes"(they all lead to loss) so any of them is as good as any other.
Im very suprrised and glad to see Humans still have something against AlphaGo, but ultimately, these kind of errors might dissapear if AlphaGo trains 6 more months. It made a tactical mistake, not a theory one.
I think there's something more subtle going on.
Lee Sedol doesn't have RAM that can be crammed with faithfully recalled gigabytes of information, and that allow exhaustive, yet precise searching of vast information spaces. The amount of short-term information Sedol can remember perfectly is very small by comparison, and doing so requires a lot of concentration and effort.
Secondly, the faculty with which Lee Sedol plays Go wasn't designed for the exclusive task of playing Go. Without having to load a different program, Sedol's brain can do many other things well.
http://i.imgur.com/ny3RhD4.png
My guess is that Sedol won because he introduced sufficient complexity through cutting points and numerous black groups (see the image). Since AlphaGo uses Value and Policy networks to determine the hot spots to analyse using Monte Carlo tree searches, by making a game rife with lots of simultaneous fights, Sedol dodged the one-two punch of Value and Policy networks combined with MCTS.
In other words, if Sedol can make over a dozen points of interest on the board, AlphaGo cannot deeply assess them all. In the image, there are at least 13 interesting moves and cuts plus up to 15 groups (depending if lone stones are considered groups by AlphaGo). I suspect that this position was far more complex than at any point during any of the three previous games.
It might also explain the meltdown of playing out an unfavourable ladder (the P10 group, as P8 is another possible move).
https://en.wikipedia.org/wiki/Go_and_mathematics#Game_tree_c...
Eventually, math wins. There will come a point where humans cannot make the game sufficiently complex to beat a domain-specific machine intelligence (such as AlphaGo).
Ke Jie is 19 years old. Lee has been a pro for nearly 20 years. Lee is not old but certainly not young. Go game prodigies seem to peak when early 20's, much like mathematicians.
My personal hypothesis for the reason why Magnus Carlsen is the youngest chess champion of all time is that with the rise of online battles and computer battles, the age where the blend of crystalized intelligence and working intelligence combines lowers.
This has obvious implications for the recent rise of machines immediately correcting human mistakes in mathematics and physics.
I'm sure there are many levels of watchdogs in this program.
> This was when things got weird. From 87 to 101 AlphaGo made a series of very bad moves.
It seems to me, that these bad moves were a direct result of AlphaGo's min-maxing tree search.
According to @demishassabis' tweet, it had had the "realisation" that it had misestimated the board situation at move 87. After that, it did a series of bad moves, but it seems to me that those moves were done precisely because it couldn't come up with any other better strategy – the min-max algorithm used traversing the play tree expects that your opponent responds the best he possibly can, so the moves were optimal in that sense.
But if you are an underdog, it doesn't suffice to play the "best" moves, because the best moves might be conservative. With that playing style, the only way you can do a comeback is to wait for your opponent to "make a mistake", that is, to stray from a series of the best moves you are able to find, and then capitalize that.
I don't think AlphaGo has the concept of betting on the opportunity of the opponent making mistakes. It always just tries to find the "best play in game" with its neural networks and tree search – in terms of maximising the probability of winning. If it doesn't find any moves that would raise the probability, it picks one that will lower it as little as possible. That's why it picks uninteresting sente moves without any strategy. It just postpones the inevitable.
If you're expecting the opponent to play the best move you can think of, expecting mistakes is simply not part of the scheme. In this situation, it would be actually profitable to exchange some "best-of-class" moves to moves that aren't that excellent, but that are confusing, hard to read and make the game longer and more convoluted. Note that this totally DOESN'T work if the opponent is better at reading than you, on average. It will make the situation worse. But I think that AlphaGo is better in reading than Lee Sedol, so it would work here. The point is to "stir" the game up, so you can unlock yourself from your suboptimal position, and enable your better-on-average reading skills to work for you.
It seems to me that the way skilful humans are playing has another evaluation function in addition to the "value" of a move – how confusing, "disturbing" or "stirring up" a move is, considering the opponent's skill. Basically, that's a thing you'd need to skilfully assess your chances to perform an OVERPLAY. And overplay may be the only way to recover if you are in a losing situation.
But after watching the summary video of AlphaGos win... I'm fascinated.
I'm sure there are thousands of resources that can teach me the rules, but HN; can you point me to a resource you recommend to get up to speed?
Alphago learns mostly from playing against itself. (And in the future they are planning to remove the crutch of starting with human generated data entirely.)
https://twitter.com/demishassabis/status/708489093676568576
ed: oops I misread!
Doesn't this imply they weren't using a single PC?
...However if it's true that it couldn't find a way out, shouldn't its probability of winning have hit ~0% much earlier? There was still a long time between when it started acting strangely and when it resigned.
I got the impression (possibly incorrectly) that AlphaGo was trying to throw curve-balls and be 'unexpected' in a way that might have 'forced' a mistake it could exploit.
That's cool to think of AlphaGo having "realizations"
Could it possibly be that both of the mistakes were bugs? Perhaps it suggested a non-sensical position such as (25.23, 13.15), and it was snapped to (19, 13) :D
Another thing to keep in mind is that AlphaGo has no "memory", so every turn it looks at the board fresh. This means if the probabilities are very close you could have it jump around a bit either due to numerical noise from floating point calculations, model errors, or just tiny differences in probability making the behavior appear erratic and quick to change "strategy".
Alphago doesn't work like that..
78 could come to symbolize humanity.
What a special moment.
It was amazing to see how Lee Sedol found the right moves to make the invasion work.
This makes me think that if the time for match was three hours instead of two, maybe a professional player will have enough time to read the board deeply enough to find the right moves.
Sure, maybe Alphago missed a winning move. But the situation was fluid both strategically and tactically, which might have been why that machine chose it's losing move and moreover, why I don't think it was just a matter of the machine failing to find a kill - especially since I think computers have been able to beat humans on pure tesuji for a while now.
Because it is actually trivial in MCTS to reevaluate moves already taken once you have a better assessment of the positions that follow it.
Anyway I'm pretty sure what happened; I have actually implemented MCTS myself and know that you can trivially update evaluations of old nodes in MCTS as the game continues and you get a better estimate of the line of play actually taken, although you wouldn't normally have a reason to do so. Basically, it can see in hindsight that a move it took was bad, but without additional calculation you wouldn't know whether the alternatives were any better, or also are worse than estimated at the time.
Also seems in-line with the way Demis was "rooting" for the human this time – they already won so now they focus on PR.
How could DeepMind prove that AlphaGo played its best in match 4 ?
Once you step on to the field, you play your best. It doesn't matter whether the opponent is a master, or a six-year-old.
EDIT> I have a friend who I've never beaten at an RTS or non-random strategy game, even though we've played hundreds of times. If I thought for a moment that he'd let me win, I'd stop playing, and I hate losing.