The future of deep learning is photonic
spectrum.ieee.org
spectrum.ieee.org
They claim quite big performance uplift over current GPUs and even better performance per watt figures.
However I haven’t seen independent benchmarks for their Envise ASIC yet.
They seem to be using a different method with an external laser to “power” the chips using fiber optic cables rather than integrating solid state lasers into the silicon. It’s also not based on LCD technology like early attempts at photonic computing which was essentially a sandwich of LCD screens and solid state detectors and amplifiers.
My gut feeling is that Envise can actually do only a few basic operations in photonics, however if these operations are sufficiently tailored for specific tasks in ML that might be good enough.
It’s not that different than what NVIDIA did with their Tensor cores they aren’t general purpose ALUs they only do a few things but they are very fast at those tasks.
ALUs and SFU/FFUs and the rest have their own defined sizes for operands, accumulators and the likes a 32bit ALU can’t work on larger operands unless it’s capable of doing some complicated tricks which are often done in the compiler rather than by the instruction decoder / scheduler in hardware and if you need more than 32 bits you are going to see a major slow down because you need to break things into smaller pieces to fit your hardware.
Same thing with matrices if you have hardware that can multiply matrices efficiently then it’s almost certainly has a defined size and it’s up to you to optimize your workload to fit in that fixed matrix size of say 256x256.
The other dimension could be smaller but will have practical limits due to thermal and/or cross-talk between elements depending on the technology used.
Optical computing is analog so the cascaded loss matters. If we assume a 1x256 array and we biased the MZMs such that we propagated a “1” to the end of the array and assume each MZM only has 0.12dB of loss; it would have 30dB of loss at the other end. So the “1” would have 1000x less power at the far end. This is only a toy problem to show how things would scale in a very simple way. In reality these devices will have much more loss per device.
But that’s fine, because I am mostly concerned with accelerating deep learning. I’m a robotics engineer and when I look at large neural networks like GPT-3 I get the sense that robotics could work well with very massive networks, even orders of magnitude larger than GPT-3 (imagine not just ingesting text and producing a stream of words, but encoding a multidimensional world state for a robot and producing a desired action based on all current and past signals).
But to put massive neural networks orders of magnitude larger than GPT-3 in to a robot requires a significant step change in the efficiency and scale of neural network compute.
So I don’t mind if their chip doesn’t do standard logic well because a regular intel chip is great at that. I just want to see significantly more powerful neural network compute. And if the Lightmatter CEO is to be believed (I don’t know), their tech could be a boon for machine learning and robotics some day.
And they're going to have the home team advantage when that happens. So that means that unless these things are as accessible to tensorflow/pytorch or whatever the heck the framework de jour is at that point (hopefully more like Jax but better), they will get no traction.
Evidence? Every single demonstrably superior CPU architecture that failed to dislodge x86 over the past 50 years. Sure, it's finally happening, but it's also 50 years later.
They did similar goofy thinking with respect to the magic transcendental unit on GPUs so it's not like they ever learned. It's not entirely about clock rate.
Lather rinse repeat for any other architecture. You can even make a network that runs best on Graphcore that way, but it won't be fun to do it. You might even get Graphcore to pay top dollar for it though as they both need some good publicity and they have lots of VC left to squander.
This also tends to be true of video games where the platform on which they were developed is the best place to play them rather than their many ports.
If you ask me the thing that makes these things even remotely interesting is the willingness from the SW side to support new HW architectures. Without that you can’t have any innovation in HW.
Or is the game plan to make them bigger & nobody cares cause no heat & fast?