So technically this version has lost every game it's ever won.
Jokes aside, it's pretty interesting to note that they were able to combine the "policy" and "value" networks. Good SO answers on the difference (https://datascience.stackexchange.com/questions/10932/differ...)
> accumulating thousands of years of human knowledge during a period of just a few days
It'd be interesting for what this would mean when things like a neural lace become a reality.
As an aside, anyone have any other links or references to others investigating learning algorithms with a 'tabula rasa' approach?