Reverse engineering a neural network's clever solution to binary addition
cprimozic.net
cprimozic.net
I can strongly recommend an excellent biography:
The Man from the Future: The Visionary Life of John von Neumann by Ananyo Bhattacharya
I knew of von Neumann because his name shows up many many times when studying computers. But I had no idea he had several equally monumental bodies of work in such a wide range of subjects.
It’s that rare biography that helps understand the history of multiple different disciplines: quantum mechanics, game theory, computer science and more.
https://www.eetimes.com/aspinity-puts-neural-networks-back-t...
>>> The following is an incomplete list of physical implementations of qubits, and the choices of basis are by convention only: [...] Qubit#Physical_implementations: https://en.wikipedia.org/wiki/Qubit#Physical_implementations
> - note the "electrons" row of the table
According to this Table on wikipedia, it's possible to use electron charge (instead of 'spin') to do Quantum Logic with Qubits.
How is that doing quantum logical computations with electron charge different from from what e.g. Cirq or Tequila do (optionally with simulated noise to simulate the Quantum Computer Engineering hardware)?
FWIU, analog and digital component qualities are not within sufficient tolerance to do precise analog computation? (Though that's probably debatable for certain applications at least, but not for general purpose computing architectures?) That is, while you can build adders out of voltage potentials quantified more specifically than 0 or 1, you might shouldn't without sufficient component spec tolerances because noise and thus error.
IMHO, Turing Tumble and Spintronics are neat analog computer games.
(Are Qubits, by Church-Turing-Deutsch, sufficient to; 1) simuluate arbitrary quantum physical systems; or 2) run quantum logical simulations as circuits with low error due to high coherence? https://en.wikipedia.org/wiki/Church%E2%80%93Turing%E2%80%93... )
>> See also: "Quantum logic gate" https://en.wikipedia.org/wiki/Quantum_logic_gate
Analog computers > Electronic analog computers aren't Electronic digital computers: https://en.wikipedia.org/wiki/Analog_computer#Electronic_ana...
Addendum: another interesting variation to try is a small transformer network, and feeding the bits sequentially as symbols. This kind of architecture could compute bignum-sized integers.
Edit: This wiki page has a good explanation of the concept: https://en.wikipedia.org/wiki/Carry-lookahead_adder
I don’t know, but cannot jump to your conclusion without much more domain knowledge.
It was just a simple feed-forward network. It can't do arbitrary amounts of repeated addition (nor repeat any other operation arbitrarily often).
Otherwise, you'd need up to 255 (or so) additions to multiply two 8 bit numbers, I think?
[1] https://en.wikipedia.org/wiki/Multiplication_algorithm#Examp...
His system is base 10000, he can add any two numbers below 5000 and come up with a single digit response in one loop. Then covert to base10 for the rest of us. Each digit in his base 10000 system has a different visual representation, like we have 0-9.
It's integer accurate so I don't think it uses sine wave approximations. Guy can remember pi for 24 hours.
I don't think the boffins or he himself got close to understanding what is going on in his head.
Maybe in the future we will think of the word "digital" as archaic and "analog" will be the moniker for advanced tech.
I.e., the network converts the binary input and output to floating point, and then it is the CPU of the host the network is running on that really does the addition in floating point.
So usually one does a bunch of FLOPs to get "emerging" behaviour that isn't doing arithmetic. But in this case, instead the network does a bunch of transforms in and out so that in a critical point, the addition executed in the network runner is used for exactly its original purpose: Addition.
And I guess saying it is "analog" is a good analogy for this..
My intuition is as follows: if I were to train this network with pencil, paper and a slide rule, I'd expect the same result. Addition (or maybe rather integration) is embedded in the abstract structure of a neural network as a computation artifact.
Sure, specifics of the substrate may "leak through" - e.g. in the pen-and-paper case, were I to round everything to first decimal space, or in the computer case, was the network implemented with 4-bit floats, I'd expect it not to converge because of loss of precision range (or maybe figure out the logic gate solution). But if the substrate can execute the mathematical model of a neural network to sufficient precision, I'd expect the same result to occur regardless of whether the network is run on paper, on a CPU, an fluid-based analog computer, or a beam of light and a clever arrangement of semi-transparent plastic plates.
I'm curious how reproducible this is. I'm also curious how networks form abstractions and if we can verify those abstractions in isolation (or have other networks verify them, like this) then the building blocks become far less opaque.
On the other hand, why we use digital hardware precisely because it's robust against the noise present in the hardware analog circuits.
I wonder, is there ever a case analog-on-digital is better to work with as an abstraction layer, or is it always easier to work with digital signals directly?
Of course, these are all just approximations of analog. Much like emulation, there are limitations, penalties, and inaccuracies that will inevitably kneecap applications when compared to a native implementation running on bare metal (though, it seems that humans lack a proper math coprocessor, so calculus could be considered no more "native" than traditional algebra).
We do sometimes see specialized hardware emerge (DSPs in the case of audio signal processing, FPUs in the case of floats, NPUs/TPUs in the case of neural networks), but these are almost always highly optimized for operations specific to analog-like data patterns (e.g.: fourier transforms) rather than true analog operations. This is probably because scalable/reliable/fast analog memory remains an unsolved problem (semiconductors are simply too useful).
Because of that, I would (hand-wavingly) expect it can be made to work for a three-bit adder (with four bits of output)
For instance instead of evaluating an exponential function in digital logic, it might be quicker and more energy-efficient to just evaluate it using a diode as an analog voltage, if the value is available as an analog voltage and if some noise is tolerable, as done with analog computers.
Which was also really cool!
> As with Thompson's FPGA exploiting some subtle physical properties of the device, Layzell found that evolved circuits could rely on external factors. For example, whilst trying to evolve an oscillator Bird and Layzell discovered that evolution was using part of the circuit for a radio antenna, and picking up emissions from the environment [18]. Layzell also found that evolved circuits were sensitive to whether or not a soldering iron was plugged in (not even switched on) in another part of the room[19]. ...
However, OP was actually referring to an experiment by Dr Adrian Thompson which was different but also sort of similar. The FPGA evolved by Thompson ended up depending on parts of the circuit that were disconnected from the main circuit but still affected its operation. It probably relied on electromagnetic properties, so was sort of a radio, but it did not rely on the clock of a nearby computer.
[damninteresting.com did a really interesting writeup about this](https://www.damninteresting.com/on-the-origin-of-circuits/)
Both were really cool, unexpected behaviors of evolved hardware systems.
Think "property rights" are required for a free market.
The human brain only consumes around 20W [1], but for numerical calculations it is massively outclassed by an ARM chip consuming a tenth of that. Conversely, digital models of neural networks need a huge power budget to get anywhere close to a brain; this estimate [2] puts training GPT-3 at about a TWh, which is about six million years' of power for a single brain.
[1] https://www.pnas.org/doi/10.1073/pnas.2107022118
[2] https://www.numenta.com/blog/2022/05/24/ai-is-harming-our-pl....
Six millions years of power for a single brain is one year of power for six million brains. The knowledge that ChatGPT contains is far higher than the knowledge a random sample of six million people have produced in a year.
With that said, I think we should apply ML and the results of the bitter lesson to things which are more intractable than searching the web or playing Go. Have you ever talked to a System's Biologist? Ask them about how anything works, and they'll start with 'oh, it's so complicated. You have X, and then Y, and then Z, and nobody knows about W. And then how they work together? Madness!"
On the other hand, I don't think if it would be larger than that of six million people selected more carefully (though I suppose you'd have to include the costs of selection process in the tally). I also imagine it wouldn't be larger than six thousand people specifically raised and taught in coordinated fashion to fulfill this role.
On the other other hand, a human brain does a lot more than learning language, ideas and conversion between one and the other. It also, simultaneously, learns a lot of video and audio processing, not to mention smell, proprioception, touch (including pain and temperature), and... a bunch of other stuff (I was actually surprised by the size of the "human" part of the Wikipedia infobox here: https://en.wikipedia.org/wiki/Template:Sensation_and_percept...). So it's probably hard to compare to specialized NN models until we learn to better classify and isolate how biological brains learn and encode all the various things they do.
See: "on the gripping hand" (-:
i'm not a physics guy but wanted to check this - a Watt is a per second measure, and a Terawatt is 1 trillion watts, so 1 TWh is 50 billion seconds of 20 Watts, which is 1585 years of power for a single brain, not 6 million.
i'm sure i got this wrong as i'm not a physics guy but where did i go wrong here?
a more neutral article (that doesn't have a clear "AI is harming our planet" agenda) estimates closer to 1404MWh to train GPT3: https://blog.scaleway.com/doing-ai-without-breaking-the-bank... i dont know either way but i'd like as best an estimate as possible since this seems an important number.
1TWh / 20 Watt brain = 50,000,000,000 (50 Billion) Hours.
50 Billion Hours / (24h * 365.25) = 5,703,855.8 Years
10^12 / 20 (power of brain) / 24 (hours in a day) / 365 (days in a year) = 5 707 762 years.
The problem is: your input dataset definitely has biases that you don't know about, and the model is going to learn spurious things that you don't want it to. It can make some indefensible decisions with high confidence and you may not know why. Some domains your model may be basically useless, and you won't know this until it happens. This often isn't good enough in industry.
To stop batshit things coming out of your system, you may have to do the opposite - use domain knowledge to break the problem the model is trying to solve down into steps you can reason about. Use this to improve your data. Use this to stop the system from doing anything you know makes no sense. This is really hard and time consuming, but IMO complex e2e models are something to be wary of.
Also, with data driven approaches, the model isn't necessarily learned in any meaningful way. If you train with certain inputs to get certain realistic looking outputs, you can build sophisticated parrots or chameleons. That's what GPT and stable diffusion are: compressed knowledge bases to get an output without a knowledge model. (No, language models are not knowledge models).
Thinking, rational or otherwise, requires causal steps. Since none of these data driven approaches have even fuzzy causal models, they require memorizing infinite universes to see if search can find a particular universe that's seen this before. That's why they're not intelligent and never will be. An intelligent animal only needs one universe and limited experiences to solve a problem because it knows the latent structure of the problem and can generate approaches. Intelligent entities, unlike these autistic Rain Man automatons, do not need to memorize every book in the library to multiply two large numbers together.
https://hn.algolia.com/?q=bitter+lesson
I'd say it has aged well and there are probably lots of new takes on it based on some of the achievements of the last 6 months.
At the same time, I don't agree with it. To simplify, the opposite of "don't over-optimize" isn't "it's too complicated to understand so never try". It is thought provoking though
> If you think about what a human, a human probably in a human’s lifetime, 70 years, processes probably about a half a billion words, maybe a billion, let’s say a billion. So when you think about it, GPT-3 has been trained on 57 billion times the number of words that a human in his or her lifetime will ever perceive.[0]
0. https://hai.stanford.edu/news/gpt-3-intelligent-directors-co...
He is claiming GPT3 was trained on 57 billion billion words. The training dataset is something like 500B tokens and not all of that is used (common crawl is processed less than once), and I'm damn near certain that it wasn't trained for a hundred million epochs. Their original paper says the largest model was trained on 300B tokens [0]
Assuming a token is a word, as we're going for orders of magnitude, you're actually looking at about a few hundred times more text. The point kind of stands, it's more, but not billions of times.
I wouldn't be surprised if I'm wrong here because they seem to be an expert but this didn't pass the sniff test and looking into it doesn't support what they're saying to me.
[0] https://arxiv.org/pdf/2005.14165.pdf appendix D
Some words are broken up into several tokens, which might explain the 300B tokens.
300B is then the "training budget" in a sense, not every dataset is used in its entirety, some are processed more than once, but each of the GPT3 sizes were trained on 300B tokens.
In fact, we are not trained on words at all. We are trained on images, sounds, tactile, and a few more types of sources.
https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mec...
One interesting thing is that this network similarly does addition using a Fourier transform.
1. How much of this outcome is due to the unusual (pseudo) periodic activation function? Seems like a lot of the DAC-like behavior is coming from the periodicity of the first layer’s output, which seems to be due to the unique activation function.
2. Would the behavior of the network change if the binary strings were encoded differently? The author encodes them as 1D arrays with 1 corresponding to 1 and 0 corresponding to -1, which is an unusual way of doing things. What if the author encoded them as literal binary arrays (i.e. 1->1, 0->0)? What about one-hot arrays (i.e. 2D arrays with 1->[1, 0] and 0->[0, 1]), which is the most common way to encode categorical data?
"While playing around with this setup, I tried re-training the network with the activation function for the first layer replaced with sin(x) and it ends up working pretty much the same way. Interestingly, the weights learned in that case are fractions of π rather than 1."
By the looks of it, any activation function that maps a positive and negative range should work. Haven't tested that myself. The 1 vs π is likely due to the peaks of the functions, Ameo at 1 and sine at π/2.
Regardless, it's not Ameo.
As for the encoding, I think it's a pretty normal way to encode binary inputs like this. Having the values be -1 and 1 is pretty common since it makes the data centered at 0 rather than 0.5 which can lead to better training results.
2. I doubt this matters at all here. For some architectures having inputs be 0 on average is useful, so the author probably just picked it as the default choice.
edit: several, but probably not very relevant https://graphics.stanford.edu/~seander/bithacks.html
(For an ANN FFT is more natural as it's a projection algorithm.)
In an NN context, given that you already have “transform with a matrix” as a primitive, probably something very much like sticking a https://en.wikipedia.org/wiki/DFT_matrix somewhere. (You are already extremely familliar with the 2-input DFT, for example: it’s the (x, y) ↦ (x+y, x−y) map.)
If you want a physical implementation of a Fourier transform, it gets a little more fun. A sibling comment already mentioned one possibility. Another is that far-field (i.e. long-distance; “Fraunhofer”) diffraction of coherent light on a semi-transparent planar screen gives you the Fourier transform of the transmissivity (i.e. transparency) of that screen[1]. That’s extremely neat and covers all textbook examples of diffraction (e.g. a finite-width slit gives a sinc for the usual reasons), but probably impractical to mention in an introductory course because the derivation is to get a gnarly general formula then apply the far-field approximation to it.
A related application is any time the “reciprocal lattice” is mentioned in solid-state physics; e.g. in X-ray crystallography, what you see on the CRT screen in the simplest case once the X-rays have passed through the sample is the (continuous) Fourier transform of (a bunch of Dirac deltas stuck at each center of) its crystal lattice[2], and that’s because it’s basically the same thing as the Fraunhofer diffraction in the previous paragraph.
Of course, the mammalian inner ear is also a spectral analyzer[3].
[1] https://en.wikipedia.org/wiki/Fourier_optics#The_far_field_a...
[2] https://en.wikipedia.org/wiki/Laue_equations
[3] https://en.wikipedia.org/wiki/Basilar_membrane#Frequency_dis...
Ooh, yes, I’d forgotten that! And I’ve actually done this experiment myself — it works impressively well when you set it up right. I even recall being able to create filters (low-pass, high-pass etc.) simply by blocking the appropriate part of the light beam and reconstituting the final image using another lens. Should have mentioned it in my comment…
H. Massalin, “Superoptimizer - A Look at the Smallest Program,” ACM SIGARCH Comput. Archit. News, pp. 122–126, 1987.
https://web.stanford.edu/class/cs343/resources/superoptimize...
The H is for "Henry".
Somewhat famously, you can plot the weight activations on a heatmap from a CNN for image processing and obtain a visual representation of the "filter" that the model has learned, which the model (conceptually) slides across the image until it matches something. For example: https://towardsdatascience.com/convolutional-neural-network-...
Many techniques don't look directly at the numbers in the model. Instead, they construct inputs to the model that attempt to trace out its behavior under various constraints. Examples include Partial Dependence, LIME, and SHAP.
Also, those "deep dream" images that were popular a couple years ago are generated by running parts of a deep NN model without running the whole thing.
> One thought that occurred to me after this investigation was the premise that the immense bleeding-edge models of today with billions of parameters might be able to be built using orders of magnitude fewer network resources by using more efficient or custom-designed architectures.
Transformer units themselves are already specialized things. Wikipedia says that GPT-3 is a standard transformer network, so I'm sure there is additional room for specialization. But that's not a new idea either, and it's often the case that a after a model is released, smaller versions tend to follow.
No, you can use transformers for vision, image generation, audio generation/recognition, etc.
They are 'specialized' in that they are for working with sequences of data, but almost everything can be nicely encoded as a sequence. In order to input images, for example, you typically split the image into blocks and then use a CNN to produce a token for each block. Then you concatenate the tokens and feed them into a transformer.
It is definitely interesting that we can do so much with a relatively small number of generic "primitive" components in these big models. but I suppose that's part of the point.
Spotted one small typo: digital to audio converter should be digital to analog.
But that is not even the real issue. The real issue that it is not analog, it is just "analog" with quotes. The neural network is executed by a digital computer, on digital hardware. So that "analog" is just the underlying float type of the of the CPU or GPU executing the network. That is still very much digital. Just on a different abstraction level.
Sure. Here is a semantics with an op-amp: https://www.tutorialspoint.com/linear_integrated_circuits_ap...
I'm pretty convinced that something equivalent to GPT could run on consumer hardware today, and the only reason it doesn't is because OpenAI has a vested interest in selling it as a service.
It's the same as Dall-E and Stable Diffusion - Dall-E makes no attempt to run on consumer hardware because it benefits OpenAI to make it so large that you must rely on someone with huge resources (i.e. them) to use it. Then some new research shows that effectively the same thing can be done on a consumer GPU.
I'm aware that there's plenty of other GPT-like models available on Huggingface, but (to my knowledge) there is nothing that reaches the same quality that can run on consumer hardware - yet.