Inside an AI 'brain' – What does machine learning look like?
graphcore.ai
graphcore.ai
I'm not the only one[1] within the research community to think this.
If you find any insights from this, I'd honestly first be surprised and then second be interested to know what insights you gleamed from it.
Background: researcher who publishes papers in deep learning.
[1]: https://twitter.com/jackclarkSF/status/834461913262157824 (thread containing a member of OpenAI who specializes in communicating complex machine learning topics to the media and a primary developer of PyTorch / member of Facebook's AI Research lab)
> I'm not the only one within the research community to think this.
I should hope not if it's indeed the consensus :) Anyhow, I agree with your sentiment.
Comparisons with organic intelligence? Check.
Vague descriptions of new technologies? Check.
Flashy, uninformative graphics? Check.
Maybe I'm being too cynical.
Yann LeCun: My least favorite description is, “It works just like the brain.” I don’t like people saying this because, while Deep Learning gets an inspiration from biology, it’s very, very far from what the brain actually does. And describing it like the brain gives a bit of the aura of magic to it, which is dangerous. It leads to hype; people claim things that are not true. AI has gone through a number of AI winters because people claimed things they couldn’t deliver."
[1] http://spectrum.ieee.org/automaton/robotics/artificial-intel...
It seems that, when educating the public about AI in order to advance the field, we must first go through this phase where people do not understand and we risk taking a step back.
I'm in love with data visualization / generative art.
- the images and false colors need to show some semblance of stability for a given network between epochs; and it needs to be robust against changing data or input structure.
- requiring visual inspection doesn't give you something you can automate with, unlike an evaluation score.
- if there is indeed a significant deviation in "MRI"-like scans between batches, its diagnostic utility ends there - it tells you nothing about what caused a change.
In the OP article's masthead, the clusters are labelled, this is AlexNet's computational graph from a Tensorflow description (depicted in full lower down the page).
On the right is "Conv1 11x11 forward [3 in, 64 out]" which suggests a Convolutional Layer with 3 inputs and 64 outputs. Forward, presumably, the direction of tensors flowing through the layer.
Alexnet's Layer 1 is Convolutional [a] :
• Images: 227x227x3
• F (receptive field size): 11
• S (stride) = 4
• Convlayer output: 55x55x96
Compare Colah's 2D Convolutional NN depiction: http://colah.github.io/posts/2014-07-Conv-Nets-Modular/or CS231n's page containing both structural and activation diagrams: http://cs231n.github.io/convolutional-networks/
In reply to your linked tweet [1], Chintala asserts the links are " compute/mem activity on their cores" and "connections between clusters are memory transfers/activity"
Does Chintala mean compute cores, Graphcore's IPUs ?
From Graphcore's page: "computational graphs are made up of vertices (think neurons) for the compute elements, connected by edges (think synapses), which describe the communication paths between vertices."
What do Graphcore's colours represent ? What is an IPU ? Is an IPU hardware like Google's TPU ?
[edit] Graphcore explains: "Our Poplar graph compiler has converted a description of the network into a computational graph of 18.7 million vertices and 115.8 million edges. This graph represents AlexNet as a highly-parallel execution plan for the IPU. The vertices of the graph represent computation processes and the edges represent communication between processes. The layers in the graph are labelled with the corresponding layers from the high level description of the network. The clearly visible clustering is the result of intensive communication between processes in each layer of the network, with lighter communication between layers."
So it is a compiled computational graph just like Tensorflow's.
I cannot fathom it further and am unsure how or if to proceed.
Counting the little dots and lines in cluster Conv1 and relating this to to 11x11 [3 in, 64 out] & Alexnet's 1st Layer could be a place to start an inquiry.
[a] http://vision.stanford.edu/teaching/cs231b_spring1415/slides...
Without any explanation of the questions you raise, this page is 99% marketing speak, and to me, next to useless.
Why graphcore is going to be any different is anybody's guess. Although, I admit the concept sounds cool on paper and the graph plots look pretty- I'd hang one on my wall for sure.
Now whether these particular professors are worth anything is a different question...
Is this just marketing mumbo-jumbo? I don't understand how a "graph processor" would look any different than a vector processor.
So basically yeah, it's marketing.
http://www.cise.ufl.edu/research/sparse/matrices/synopsis/
I think Trefethen too has nice visualization like this (or maybe that was the spectral thing).