Optical Transformers
arxiv.org
arxiv.org
Running the math on a machine with 8x A100 (enough to run today's LLMs), that would be 300w * 8gpus / 100 = 24w.
This is within striking distance of IOT and personal devices. I'm trying to imagine what a world would look like where generative text models are commodetised to the point where you can either generate text locally on your phone, or generate GBs of text in the cloud.
I have to admit it's very hard to make any sort of accurate prediction.
Maybe I interpreted that incorrectly but I thought it's saying a 100x advantage for current large Transformer models, and 8000x advantage for future quadrillion-parameter models? I didn't include those because I suppose that size of model is quite a few years away. Admittedly this is only based on the abstract...
If most devices are replaced within a year or two then you get a pretty good cadence for updating your Siri model (and even more incentive for users to upgrade hardware).
It's possible to have a temporary lead because of constant factors, but as long as an electronic circuit has to expend a unit of energy per MAC, you'll always be able to specify a model big enough that an optical network will beat it.
> We conclude that with well-engineered, large-scale optical hardware, it may be possible to achieve a 100× energy-efficiency advantage
Emphasis on may.
2) in the real world, constant factors matter (as you allude to). For example if an ASIC gets a 1000x speedup (optimistic; we saw this for BTC) it might be the better choice for this generation, but start to lose next gen and beyond. If an ASIC only gets 100x or lower then it’s not favorable this gen.
So sure, this tech might win in the long term, but I wasn’t making any categorical claims, just noting that there are multiple horses we need to track.
It would be quite foolish to dismiss custom silicon solutions based on this paper.
IIRC one reason we don't already have fully 3D chips is because of the heat dissipation. Reducing 2400 W to 24 W means the heat is much more tractable, which means it can be closer to volumetric than planar.
Consider a 1cm*1cm*1mm chip; with 1μ^3 elements, 1e11 per chip; with (10nm)^3 elements limited to one layer because of heat, 1e12.
Yes this is still a factor of x10, and chips are a few layers because while heat is a problem it's not a total blocker, but it's still much less than the 100^3 ratio a simple scale-up would result in.
Dead internet theory for one. Scalable spear phishing and scams. Scalable automated offensive hacking. SEO far worse than anything possible today. Mass manipulation campaigns.
Social interaction would also be strange. Every messenger and dating app able to automatically reply and suggest sophisticated messages.
You can get a hundred dollar SSD that has a read speed around 7 gigabytes per second using just over five watts. That will fill up 8x80GB in a minute and a half of load time. If your energy budget is 24 watts then install four and make it 20-25 seconds.
As far as cost, I don't know what the proposed chip would be, but $400 is 0.5% of that pile of GPUs and SSDs will only get cheaper.
This ties really nicely to the photonic DSP as convolution is fundamentally composition of convolutive systems. This is a generalized convolution, not the standard one.
The learning mechanism of transformer models was poorly understood however it turns out that a transformer is like a circuit with a feedback.
I argue that autodiff can be replaced with what I call Hopf coherence which happens within the single layer as opposed to across the whole graph.
Furthermore, if we view transformers as Hopf algebras, one can bring convolutional models, diffusion models and transformers under a single umbrella.
I'm working on a next gen Hopf algebra based machine learning framework. The beautil part is that it ties nicely to PL aspect as well.
Join my discord if you want to discuss this further https://discord.gg/mr9TAhpyBW
I wish I was smart enough to know what a Hopf algebra was or how it worked, because this sounds awesome.
Look at the diagram. Do you see the path going through the middle? And the paths at the top and bottom (they are generalized convolutions)? Well a Hopf algebra "learns" by updating it's internal state in order to enforce an invariance between the middle path and the top and bottom paths.
Allow me to restate it, it's an algebra that "learns". Reading about Hopf algebras is tripper than dropping acid.
I first met this with a floating-point systems co-processor on a Dec-10. It was highly problem specific, only a single DECUS tape FORTRAN compiler talked to it (IIRC) and you had this giant freezer sized box (it was lime green, unlike the blue DEC-10 livery) sat next to the CPU doing... mostly nothing. It was basically idle almost all the time. Occupying floorspace, consuming power, making the DEC-10 more complex to operate.
Same-same with 3DES processing cards, on-card ethernet TCP, you-name-it -these things absolutely do improve the world, but at a cost of complexity.
So the idea of designing optical interference/coherence/diffraction "engines" to bolt onto the side of a beautiful ISA, and winding up with something morally like a GPU which has to do complex TCAM like memory, wierd protocols on the bus, all to get what? some context specific speed up of .. text processing?
If this winds up down in the VLSI level on-chip, as an alternate compute element inside the ALU/ISA space, with its own bus path, and simply integrates more tightly into the ISA I'd be fine.
Perhaps thats where it's going.
Well, the beautiful (not so beautiful these days) ISA is always there for you, while we grab a library to use a suitable specialized processor and run our shit 1000x faster for 1/100 the power.
I feel like this about the TPU and other chips bundled into phones too. I wonder if we're going to wind up with "photoshop specific" bolt on architectures.
I'm not trying to be combatitive, but it seems completely arbitrary where one would draw a line and taken to the extreme one is essentially left with a CPU and possibly some RAM that might be some ideal computing machine, but can't even talk to the outside world.
So the whole 'hyperthreading is cool, 2x the CPU' thing: it turns out that as long as your e.g. integer only, true. If you invoke FP, not true: they stall.
L1 cache in some instances is now big enough some people's code never leaves it.
Nobody who lives above a compiler should ever have to worry about this. I live above scripts. I don't really worry about this. I just hark back to when I ran these boxes, on raised floors, remembering how .. distasteful adjunct computing equipment was. It was completely asynchronous, required it's own resets, interfered with cabling, had effects on the bus (unibus!) you didn't entirely understand, they made life "harder" for operations.
Did they speed things up? Sure!
CPU? Ooh you lucky, lucky sod. Back in my day, you wanted an OR gate you had to string some NANDS together.
I have worked on systems doing HSM like activity which depended on vendor specific TPM. They're a pain.
From that perspective, speedup is less about hyperefficient design of a brand new thing, and more about recognizing large portions of existing workload that are amenable to acceleration.
More successful accelerators have been borne from the latter approach (superscalar/hyperthreading, GPUs once OpenGL/DirectX dominated, fixed function video decode hardware) than the former (Itanium/VLIW)*.
Or in simpler form, never start a value proposition with "First, rewrite all your software..."
* ML is an odd duck, as it's somewhat co-evolving with its own accelerators?
I am not in this deep. I am basically howling at the moon, at this point.
If you're looking at deep learning, the end game is probably hardware a bit like the brain: a hundred billion of specialized, low-clocked coprocessors crammed in a small space with only local communication to their neighbours.
CUDA and subsequent libs arguably provided this and caused a pretty massive acceleration in the amount of GPU-utilising code that got written. Now, CUDA might still be pretty awful to write, but if general-propose-GPU API’s continue to evolve and refine, we might be able to get to a point where something like an LLVM compiler can auto generate efficient GPU code for matrix ops, much like how it can auto-vectorise certain patterns now.
"...By the way, for the sake of consistency, AOBJN should have been called AOBJL and AOBJP should have been called AOBJGE. However, they weren't. ..."
With Moore’s law dominating any ideas about software optimization, the mantra of “wait until the CPU is twice as fast in two years” has dominated the industry for long enough.
Not everything needs to fit in an idealistic CPU and we should allow for more complex computational machinery.
Let’s allow for new hardware, new programming paradigms, new ways of thinking.
FPGA's and ASIC were tuned to do highly specific tasks, and then wound up being much the same: prove it in an FPGA, make an ASIC, then see it move on-die as the VLSI matures.
Now Moores law has bottomed out, specific purpose processing asynchronously makes more sense.
I just don't like it, the same way angry old men growl at the moon at night.
The 8-bit era and before relied heavily on co-processors. Sometimes they were the same or similar to the main CPU - e.g. the Commodore 64 floppy drive had a CPU as powerful as the main CPU of the machine, and you could run code on it - but it took special effort to target them. The most successful 16-bit home computers had co-processors and special purpose hw (e.g. the Amiga had it's copper and blitter, but some models also had a 6502-compatible core controlling the keyboard, and many SCSI controllers at the time had full CPU's - mine had a Z80).
The PC got cheap by relying on the CPU to be able to control most things in an age where special purpose hardware was expensive due to low volumes (the Amiga "only" sold about 5 million across multiple models), but it didn't take long before a typical PC started getting micro-controllers and full on CPUs everywhere (e.g. in your hard drive [1]) because the CPU speed as mostly not gotten fast enough.
Just for a short window of maybe a decade it got fast enough to outpace the cost of low volume special purpose HW.
That said... Can anyone parse what's being done here? How/why is it more efficient? Is this essentially the Analog Computer of these giant transformer models?
Finally, how can I build one with fiber optic cable and leds?
https://en.wikipedia.org/wiki/Spatial_light_modulator
I've seen a lab bench prototype of a different implementation, there are a lot of engineering problems to solve but as the paper points out the potential payoff is big.
Edit: The other key point is that one of the expensive components in transformers is effectively a giant matrix multiplication which implies many, many individual multiplications.
To get some intuition about the promise imagine being able to implement the weights of a layer (fully connected layer is essentially vector matrix multiplication) as a 2D hologram and the compute as pushing an image from an OLED display through to an image sensor. Multiplication as attenuation in the hologram, summation is just the accumulation of charge on the sensor side. Everything happens all at once and the number of photons required is potentially very small. An actual working implementation would be both more clever and more practical. The potential to do every multiply in parallel for almost no energy is so attractive that I expect people to chip away at this problem for the foreseeable future.
You also don't need very high precision: many models perform well even using 8-bit floats. This allows to ditch the whole digital approach and implement analog circuits, which, while sometimes less precise, are massively simpler and more energy-efficient.
So they built a mostly-analog device specialized for running ML models, and used optics instead of electronics, which allows to make certain things much faster, on top of the simplification.
There are known attempts to use electronic approaches, such as using flash-like structures that store charge as an analog value to store model weighs, and to do addition and multiplication right inside the cells.
"On-chip optical matrix-vector multiplier for parallel computation" (2013)
https://spie.org/news/4932-on-chip-optical-matrix-vector-mul...
Figure 1 has a nice visual of the vector-matrix computation in terms of the diode laser array, the multiplexer, the "microring modulator matrix" and the resulting output, the new (output) vector detection system.
> "We have designed and fabricated a prototype of a system capable of performing a multiplication of a M × N matrix A by a N × 1 vector B to give a M × 1 vector C. The mathematical procedure of MVM can be split into multiplications and additions, which is reflected in our design. Figure 1 shows a schematic of the architecture we propose. The elements of B are represented by the power of N modulated optical signals with N different wavelengths (λ1, λ2, …, λN), generated by N modulated laser diodes, either alone or together with N Mach-Zehnder modulators. These signals are multiplexed, passed through a common waveguide, and then projected onto M rows of the modulator matrix by a 1 × M optical splitter. Each element aij of matrix A is represented physically by the transmissivity of the microring modulator located in the ith row and the jth column of the modulator matrix. Each modulator in any one row only manipulates an optical signal with a specific wavelength."
That was 10 years ago, no idea what current state-of-the-art is.
A custom controller sets flash bits to arbitrary values between 0 and Vcc, normal flash reading results in product calculation.
The main reason for this is I want something I can hum a song to and it tells me what the song is. We have Shazam, but AFAIK that works almost entirely using fourier transforms / audio signals in frequency space, and so doesn't work well if it's not an actual snippet of the song. But it would be cool if we could train an audio model like we've trained language models to "fill in the gaps" for me. Potentially there could be some interesting applications where you sing an entirely new song and the model could "fill in the gaps" with instruments, etc.
Would this sort of analog-digital circuitry or something similar possibly work to build a pitch shifter with high enough resolution to effectively allow full audio range pitch shifting of source audio without warbling? Of all the things I've tried, the SBLive! was actually the best, holding semi-stable 2 full steps down on my guitar. Everything else was fairly nasty after a full step.
Choosing positioning like "NNs are similar to the brain" or even calling the field "machine learning" makes it harder to speak objectively about the research we perform, makes it hard to understand the sort of explanatory power and limitations of the models we fit, and makes it harder for users to understand how the system works. It's like how quantum mechanics researchers have to deal with readers who misunderstand what linear observable operators are and conflate it with their own ideas about conscious observers or use it as evidence of God's presence or whatever.
Activation functions in NN map to action potentials/postsynaptic potentials, weights map to synaptic strengths/connection patterns between neurons, dropout to synaptic pruning/apoptosis, convolutional layers from the receptive fields of our visual cortex.
Certainly our brains are more complicated. Temporal spike timing, dendritic signals, synaptic plasticity, compute of synapse themselves, lots more I can't recall, none of these concepts are currently modelled in our NN architectures.
NNs start from the simple premise of interconnected neurons. As the field develops and we find what works or what doesn't, perhaps we will take more bits of inspiration from the brain.
A Perceptron model is just a dot product: one multiply+add, summed to one output scalar value. Modern nomenclature might consider Perceptrons to just be "one channel of a single layer," or like "one fully-connected layer with a single output." In fact, that's where the old-fashioned name for neural networks came from: "Multi-layer Perceptrons (MLPs)" are a bunch of these arranged on top of each other. The older 1943 work showed that in some sense lots of possible arrangements are roughly equivalent to each other so it made sense for Rosenblatt and friends to start from the simplest model with the understanding that the field would build from there. (Our feedforward model ancestry is just one branch of the tree rooted in this work; there's a completely separate geneology for other possible arrangements of modeling components, like recurrent Boltzmann machines that simply haven't been so well studied/understood... :-)
The generality of Perceptrons made room for a lot of interpretive flexibility. Right from the start, news reporters saw the theory, attended academic conferences, and started writing about it. The NYTimes had an article in 1958 that spoke of the Perceptron as "the embryo of an electronic computer that [the Navy] expects will be able to walk, talk, see, write, reproduce itself and be conscious of its existence."
In fact there was a very famous kerfuffle between Rosenblatt (Perceptron authors) and Minsky & Papert (another AI researchers). The latter two published a spicy book in 1969 saying that a single Perceptron wasn't general enough because linear models can't even learn an XOR function. Everyone who read this book misinterpreted it as a takedown of Perceptrons; this book inadvertantly dried up interest, choked funding across the board, and ultimately caused the first AI winter which didn't get resolved until well into the '80s when these ideas started to be revitalized. You can read about the spilled tea here, it's quite fascinating: https://doi.org/10.1177%2F030631296026003005
[1]: https://en.wikipedia.org/wiki/Optics [2]: https://hackage.haskell.org/package/optics-0.4.2/docs/Optics...
And then the post-singularity open source release, GPT-ℵ₁
I wish more researchers made that kind of commitment when they publish.
That said, photonic gates for digital circuits hold some promise for power efficiency and speed in their own way. For many of the parameters which apply to crypto mining, it's not so clear whether photonic digital circuits will significantly outperform the best near-future electronic transistors though.
(Source: I worked on a photonic crypto circuit design based on 4-wave (non-linear) mixing. Non-linearity is required for digital logic. Even though 4-wave mixing is energy efficient, the design was less efficient and performant than I'd initially hoped from the high spatial parallelism and basic THz figures. The principles still hold a lot of promise.)
Now for some specific operations optical signal processing might still make sense, for example if you can take advantage of the bandwidth and inherent parallism.