Why Are Eight Bits Enough for Deep Neural Networks?
petewarden.com
petewarden.com
For me, the place we'll eventually end up is obviously custom deep learning / evaluation chips that perform analogue operations using transistors in their linear regime (like how op-amps work). These chips would be programmed merely to express the tensor operation graph, essentially analogue tensor FPGAs.
This should bring multiple order-of-magnitude reductions in power consumption and increases in evaluation speed. And when you don't have a clock there might also be interesting ways of dealing with time in which you don't discretize and unroll, like one currently does with GRUs or LSTMs.
This does not make sense to me. Can you explain?
I think there might be misunderstanding of how analog computing is used to build a neural network. First, a weight is stored as some analog physical property, typically as charge on a floating gate, or on a capacitor in a DRAM type cell. Second, the multiplication operation is performed by modulating the analog input signal going through the floating gate transistor by the charge on the floating gate (weight). Third, the summation is done via simple summation of the currents. Finally, activation function is performed by an opamp.
Regarding power consumption: 1. A digital computer needs a thousand of transistors to perform multiplication, analog circuit can do it with a single one. 2. Analog NN stores parameters (weights) locally, right where they are needed to perform computation. Digital NN will need lots of memory transfers to bring weights from RAM to ALU, and to store intermediate results.
That's why a properly implemented analog NN will always consume much less power.
That's interesting. What would the circuit be?
> Digital NN will need lots of memory transfers to bring weights from RAM to ALU, and to store intermediate results.
That's not necessarily the case. Cellular neural networks were proposed long ago, for example, and they're digital -- how multiplication happens is independent from the data flow architecture.
> That's why a properly implemented analog NN will always consume much less power.
How do you know that the I^2 cost of operating in the linear regime isn't excessive? I'm totally ignorant on the matter -- I'd love to see a ballpark calculation to understand why it isn't important.
Cellular neural networks were proposed long ago, for example, and they're digital
What is so inherently digital about cellular networks? Can you provide a link to an implementation of a cellular net in digital hardware? How the weights are stored? Where the multiplication happens?
I understood the reasoning to be that to increase the range of accurately representable values in a circuit, you either need to increase the voltage or current used in an analog circuit (to achieve a certain accuracy versus a noise baseline), or devote more bits in a digital circuit. The first gives a linear dependence (or quadratic for I^2 losses) of power on range, the second logarithmic.
To give a bit of a dramatic illustration, if you circuit has on the order of 1 nV of thermal noise and you wanted to do the linear analog equivalent of 64bit integer arithmetic, you would need a signal on the order of 10,000,000,000 V to have enough precision. In fact, in terms of power consumption it's even worse. If the 1 nV signal consumes something like 1 pW, you would need something like the total power output of the Sun (on the order of 10^26 W) -- a bit of an expensive multiplication, no :) ? That's how crazy it is!
Again, if you can get away with less than 8 bits of precision and imperfect linearity the picture changes, but I wouldn't declare it superior a priori without looking at the numbers.
But yes, I understand your point. Both analog and digital implementations have their strengths and weaknesses. If you value power over precision, go with analog. If the opposite - go with digital.
The reason NNs don't exhibit strong error propagation is because of the non-linearities between linear layers that perform operations analogous to threshold/majority voting or the like, which have error correction properties.
Be careful with jumping to conclusions: I never even cited ReLUs or Sigmoids in my post! I don't have any opinion on which non-linearity is better, I only know both are dramatic non-linearities. My claims were about linear circuits. You should use whatever nonlinear element works best in your Neural Network, of course (and I've heard ReLUs have good advantages).
That's true, I feel stupid for not having thought of that!
I'm not an electrical engineer, but with the FETs that modern Intel chips are using, what fraction of their power consumption comes from parasitic gate capacitance, versus other losses?
And if you operated in the linear region, what's the ballpark steady-state I_SD current you'd need on one FET to drive the gate of the next FET?
I think that's what this comes down to: if gate capacitance dominates other losses even in the linear regime, you still win by not having a clock and lots of digital transitions.
You could even imagine exploiting that: apply 'slow' augmentations of the input data that get you the equivalent of a bunch of iterations on a single example batch, while incurring a much smaller fraction of that initial cost because activations aren't going to change nearly as much as switching to a whole new example batch.
Power is mainly lost via leakage (the smaller the transistor, the more it leaks), and via interconnect capacitance, which dominates all other capacitances in modern circuits.
Of course, but the 'ideal' FET has zero gate capacitance, despite that being the way they work.
> Power is mainly lost via leakage (the smaller the transistor, the more it leaks), and via interconnect capacitance, which dominates all other capacitances in modern circuits.
Interconnect meaning things like the buses? There's no reason to want a von Neumann architecture for an analog chip. If that leaves leakage, I suppose an analog chip would be the beneficiary of needing a lot fewer transistors per op.
I don't understand this statement. What do you mean? A FET is a capacitor (gate to channel). If a gate has no capacitance, you have no transistor.
Interconnect means wire. This has nothing to do with von Neumann architecture. If you have wires in your circuit, then you have wire capacitance. As transistors get smaller, that capacitance starts to dominate internal transistor capacitances.
It's still not clear whether the future of AI will even involve neural networks at all. Intuitively, they seem so inefficient.
For example, dropout, which is almost ubiquitous for deep learning, basically makes activations 'wrong' 50% of the time during training.
Dropout is not the same as random noise. By using dropout you eliminate some neurons from making contribution. As a result, you effectively train many smaller nets, each one adjusting its available weights to perform the same task. During testing, there's no noise - all neurons are back in business and contributing.
I was speaking loosely -- dropout is multiplicative Bernoulli noise on the hidden layers.
> That's not true. NNs don't like noise, there have been a lot of research done about effect of noise on NNs in the 90s. Random noise over a certain threshold will progressively degrade the performance of NNs, and below the threshold will have no effect.
I'd argue that dropout (and its predecessor in denoising autoencoders) are perfectly valid to see as noise, albeit multiplicative.
What would change with temperature that would require retraining? Are you saying the output of an op could depend sensitively on temperature, or that higher temperatures would increase things like thermal or shot noise? Why would the latter require retraining?
it's interesting, NN degrade at about 6bit, and that's mostly because the transfer function become stable and the training gets stuck more often in local minimums.
we built a training methodology in two step, first you trained them in 16bit precision, finding the absolute minimum, then retrain them with 6bit precision, and the NN basically learned to cope with the precision loss on its own.
funny part is, the less bit you have, the more robust the network became, because error correcting became a normal part of its transfer function.
we couldn't make the network solution converge on 4bit however. we tried using different transfer function, but then ran out of time before getting meaningful results (Each function needs it's own back propagation adjustment and things like that take time, I'm not a mathematician :D)
Stochastic rounding can fix this. You round each step with the probability so it's expected value is the same. Usually it will round down to 0, but sometimes it will round up to 1.
Relevant paper, using stochastic rounding. Without it the results get worse and worse before you even get to 8 bits. With stochastic rounding, there is no performance degradation. You could probably even reduce the bits even further. I think it may even be possible to get it down to 1 or 2 bits: http://arxiv.org/abs/1502.02551
The relevant graph: https://i.imgur.com/cOZ4fn3.jpg
This isn't true, modern SIMD instruction sets have tons of operations for smaller fixed point numbers, as used heavily in video codecs. Unless the author meant some sort of weird 8 bit float?
Also, there is no strong need for "zero".
The bits per node just determine the 'resolution' of your individual nodes; while the network as a whole determines how many states can be represented.