many people seem to discount this thread because "Karpathy is one of the most successful PhD students in his field"
perhaps instead of discounting his experience... it would be better to take his advice
41 karma · joined July 19, 2013
perhaps instead of discounting his experience... it would be better to take his advice
https://iamtrask.github.io/2016/02/25/deepminds-neural-stack...
For every node in every other layer, I colocate the edge on the same machine. In this way, when a group of, say, 10 nodes in layer 1 are each sending a weighted message to a single node in layer 2... they can pre-combine their messages (weighted sum) and send only that value over the network. This happens for every node in the second layer, reducing network i/o (this is the first optimization).