Advances in semiconductors are feeding the AI boom
spectrum.ieee.org
spectrum.ieee.org
My impression is that the main obstacle to achieving a comparable volumetric density is that we haven't cracked 3d stacking of integrated circuits yet. Very exciting to see TSMC making inroads here:
> Recent advances have shown HBM test structures with 12 layers of chips stacked using hybrid bonding, a copper-to-copper connection with a higher density than solder bumps can provide. Bonded at low temperature on top of a larger base logic chip, this memory system has a total thickness of just 600 µm...We’ll need to link all these chiplets together in a 3D stack, but fortunately, industry has been able to rapidly scale down the pitch of vertical interconnects, increasing the density of connections. And there is plenty of room for more. We see no reason why the interconnect density can’t grow by an order of magnitude, and even beyond.
It's hard to imagine not getting unbelievable results when in 10-30 years we have GPUs with a comparable number of transistors to brain synapses that support computation speed 10,000x faster than the brain. What a thing to witness!
Yes. If we could stack transistors in the Z dimension as closely as we do in X and Y, we'd easily exceed the brains density.
I think they're equivalent to a parameter AND the multiplier. Or in analog terms they'd just be a resistor whose value can be changed. Digital stuff is not a good fit for this.
For what it's worth, that's actually a thing (ReRAM/memristors), but I think it got put on the back burner because it requires novel materials and nobody figured out how to cost-effectively scale up the fabrication versus scaling up flash memory. I saw some mention recently that advances in perovskite materials (a big deal lately due to potential solar applications) might revive the concept.
Even we have 10 times higher clock speed, a trillion transistor GPU is still very slow comparatively speaking ?
Surely memory has to come in somewhere?
One important reason for the difference in current is that transistors need to reliably switch between two reliably distinguishable states, which requires a comparatively high current, whereas synapses are very analog in nature. It may not be possible to reach the brain’s efficiency with deterministic binary logic.
We know things like sleep, hunger, fear, and stress all impact how we think, yet people want to still build this mental model that synapses are just dot products that either reach an activation threshold or don't.
I would imagine the baseline assumption of your thinking is that things like sleep and emotions are a 'bug' in terms of cognition (or at the very least, 'prompts' that are optional).
Said differently, the assumption is that with the right engineer, you could reach human-parity cognition with a model that doesn't sleep or feel emotions (after all what's the point of an LLM if it gets tired and doesn't want to answer your questions sometimes? Or even worse knowingly deceives you because it is mad at you or prejudiced against you).
The problem with that assumption is that as far as we can tell, every being with even the slightest amount of cognition sleeps in some form and has something akin to emotional states. As far as we can prove, sleep and emotions are necessary preconditions to cognition.
A worldview where the 'good' parts of the brain (reasoning and logic) are replicated in LLM but the 'bad' parts (sleep, hunger, emotions, etc.) are not is likely an incomplete model.
No airplanes do not sleep. That's part of why their flying is fundamentally different than birds'.
You'll likely also notice that birds flap their wings while planes use jet engines and fixed wings.
My entire point is that it is foolish to imagine airplanes as mechanical birds, since they are in fact completely different and require their own mental models to understand.
This is analogous to LLMs. They do something completely different than what our brains do and require their own mental models in order to understand them completely.
Ornithopters are designed by humans who sleep - the complex computers needed to make them work replicate things humans told them to do, right?
It is a very incomplete model of an ornithopter to not include the human.
Yes, sleep is in fact a prerequisite to planes flying. We have very strict laws about it actually. Most planes are only able to fly because a human (who does sleep) is piloting it.
The drones and other vehicles that can fly without pilots were still programmed by a person (who also needed sleep) FWIW.
That seems implausible. Apple's M2 has 20 billion transistors and draws 15 watts at full power [1]. Even assuming that 90% of those transistors are for cache and not logic, that would still be 2 billion logic transistors * 1 milliampere = 2 million amperes at full power. That would imply a voltage of 7.5 microvolts, which is far too low for silicon transistors.
[1] https://www.anandtech.com/show/17431/apple-announces-m2-soc-...
? you sure about that? in a single transistor? over what time period, more than nanoseconds? milliamps is huge, and there are millions of transistors on a single chip these days, and with voltage drops of ... 3V? .7V? you're talking major power. FETs should be operating on field more than flow, though there is some capacitive charge/discharge.
Also the transistors are working at 1V or lower, but as you say they are FETs and don't have the same Vbe drop as a BJT.
Just add a factor 2^D transistors for each original "brain transistor" and re-run your hardware. Hope field effects don't count, and cross your fingers that neurons are idempotent!
Easy! /s
Modelling an analog system in digital will always have a combinatorial curse of dimensionality. Modelling a biological system is so insanely complex I can't even begin to think about it.
On the complexity, AFAIK a synapse is way more complex than a transistor. Larger too, if you include its share of the neuron's volume. And yes, the count difference is due to the 3D packing.
The current systems simulate stuff through computation using switches.
It’s like real sand vs sand simulator in the browser. One spins your fans and drains your battery when showing you 1000 particles acting like sand, the other just obeys laws of physics locally per particle and can do millions of particles much more accurately and with very slight increase in temperature.
Of course, analog computations are much less controllable but in this case that’s not a deal breaker.
Not even remotely comparable
* its unlikely synapses are binary. Candidly they probably serve more than one purpose.
* transistor count is a bad proxy for other reasons. A pipeline to do floats are not going to be useful for fetch from memory. "Where" the density lies is important.
* Power: On this front transistors are a joke.
* The brain is clockless, and analog... frequency is an interesting metric
Binary systems are going to be bad at simulating complex processes. LLM's are a simulation of intelligence, like forecasting is a simulation of weather. Lorenz shows us why simulation of weather has limits, there isnt some magical math that will change those rules for ML to make the leap to "AGI"
i know of at least one startup working with that concept[1].
Im sure there are others.
https://www.sciencedirect.com/science/article/pii/S089662732...
So we are going to need a lot of computational power to approximate what’s going on in an entire human brain.
“Living things” are not static designs off the drafters table. They’ll never intelligent from their own curiosity, but from ours and the rules we embed. No matter how hard we push puerile hallucinations embedded by Star Trek. It’s still a computer and human agency does not have to bend to it.
Partly, but also because the brain has an asynchronous data-flow design, while the GPU is synchronous, and as you say clocked at a very high frequency.
In a clocked design the clock signal needs to be routed to every element on the chip which requires a lot of power, the more so the higher the frequency is. It's a bit like the amount of energy used doing "battle ropes" at the gym. The heavier the ropes (cf more gates the clock is connected to), the more power it takes to move them, and the faster you want to move them (cf faster clock frequency) the more power it takes.
In a data-flow design, like the brain, there is no clock. Each neuron fires, or not, independent of what other neurons are doing, based on their own individual inputs. If the inputs are changing (i.e. receiving signal spikes from attached neurons), then at some threshold of spike accumulation the neuron will fire (expending energy). If the inputs are not changing, or at a level below threshold, then the neuron will not fire.
To consider the difference, imagine our visual cortex if we're looking at a seagull flying across a blue sky. The seagull represents a tiny part of the visual field, and is the only part that is moving/changing, so there are only a few neurons who's inputs are changing and which themselves will therefore fire and expend energy. The blue sky comprising the rest of the visual field is not changing and we therefore don't expend any energy reprocessing it over and over.
In contrast, if you fed a video (frame by frame) of that same visual scene into a CNN being processed on a GPU, then it does not distinguish between what is changing or not, so 95% of the energy processing each frame will be wasted, and this will be repeated frame by frame as long as we're looking at that scene!
Clock only needs to be distributed to sequential components like flip flops or SRAMs. The number of clock distribution wire-millimeters in typical chip is dwarfed by the number of data wire-millimeters, and if a neural network is well trained and quantized activations should be random, so number of transitions per clock should be 0.5 (as opposed to 1 for clock wires), meaning that power can't be dominated by clock. The flops that prevent clock skew are a small % of area, so I don't think those can tip the scales either. On the other hand, in asynchronous digital logic you need to have valid bit calculation on every single piece of logic, which seems like a pretty huge overhead to me.
There's more promise in analog chip designs, such as here:
https://spectrum.ieee.org/low-power-ai-spiking-neural-net
Or otherwise smarter architectures (software only or S/W+H/W) that design out the unnecessary calculations.
It's interesting to note how extraordinarily wasteful transformer-based LLMs are too. The transformer was designed part inspired by linguistics and part based on the parallel hardware (GPU's etc) available to run it on. Language mostly has only local sentence structure dependencies, yet transformer's self-attention mechanism has every word in a sentence paying attention to every other word (to some learned degree)! Turns out it's better to be dumb and fast than smart, although I expect future architectures will be much more efficient.
Or by a more static design. A GPU can't do a thing without all the weights and shaders. There are benefits of this, you can easily swap one model for another. Human mind from the other hand is not reprogrammable. It can learn new tricks, but you cannot extract a firmware from one person and to upload it to another person.
Just imagine if every logical neuron of AI was a real thing, with physical connections to other neurons as inputs. No more need to have a high throughput memory, no more need to have compute units with gigaherz frequency.
But, yes, the brain continues to be a surprising machine and ML accomplishements are amazing for that machine.
Brain is a spiking network with mutable connectivity, mostly asynchronous. Only the active path is spending energy at a single moment in time, and "compute" is tightly coupled with memory to the point of being indistinguishable. No need to move data anywhere.
In contrast, GPUs/TPUs are clocked and run fully connected networks, they have to iterate over humongous data arrays every time. Memory is decoupled from compute due to the semiconductor process differences between the two. As a result, they waste a huge amount of energy just moving data back and forth.
Fundamental advancements in SNNs are also required, it's not just about the transistors.
You can not reasonably compare this to model parameters or transistors at all.
But we know very little about how biological brains actually work and very few connectomes have been fully mapped as of yet. We still cannot fully explain how C. elegans with 302 neurons "thinks".
Also every neuron contains the entire genome, which is hundreds of millions to billions of base pairs depending on the animal and we also don't really know how important it is to the function of the brain.
So the "state" of biological brains is potentially utterly gigantic even for simple animals.
NVIDIA just announced Blackwell which gets to 208bn transistors on a chip by stitching two dies together into a single GPU. https://www.nvidia.com/en-us/data-center/technologies/blackw...
They’re sticking two of them in a board with a Grace CPU in between, then linking 36 of those those boards together in racks with NVLink switches that offer “130TB/s of GPU bandwidth in one 72-GPU NVLink domain”.
In terms of marketing, NVidia calls one of those racks a GB200 NVL72 “super GPU”.
So on one level NVIDIA would say they already have a GPU with ‘trillions’ of transistors.
Think super-[cross-fire/sli].
Economics will probably forbid that. This is a virtual limit which factors in the previous physical limits... in theory.
Arguing about what is the size limit to consider something a GPU or not is a bit like bikeshedding.
As to why wouldn't a supercomputer be considered for this? Because it's not a single chip.
One author is chairman of TSMC.
The other author is Chief Scientist of TSMC.
This is important to note because they clearly know some stuff and we should listen.
We need that ONE paper on analogue to end this quest of trillions and counting transistors.
Like people weren't trying to make computers out of bigger and bigger tubes before the transistor, they were trying to make them out of smaller and smaller ones.
Like how the transistor made the big and hot vacuum tubes obsolete, maybe we’ll see some analog breakthrough do the same thing to transistors, at least for AI.
I doubt there is a world where we use analog for general purpose computing, but it seems perfect for messy, probabilistic processes like thinking.
The software is a different story. Sure, the brain does all sorts of things that aren't necessary for $TASK, but we aren't necessarily going to be able to correctly identify which are which. Is your inner experience of your arm motion needed to fully parse the meaning in "raise a glass to toast the bride and groom", or respond meaningfully to someone who says that? Or perhaps it doesn't really matter - language is already a decent tool for bridging disjoint creature realities, maybe it'll stretch to synthetic consciousness too.
I have some theories that this isn't necessary. 1.) Just because the brain is a general-purpose machine great at doing lots of things, doesn't mean it's great at each of those things. Like when two people are playing catch, and one of them sees the first fragments of a parabola and estimates where the ball is going to land- a computer can calculate that way more efficiently than a mind, despite the fact that both are quick enough to get the job done. 2.) While the brain is great at, say, putting names to faces... a good CV machine can do the job almost as well, and can annotate a video stream in real-time.
Combining 1.) the fact that some problems are much simpler to solve with classical algorithms instead of neural networking, and 2.) that many brain tasks can be farmed out to a coprocessor/service, my hypothesis is that the number of neurons/resources required to do the "secret sauce" part of agi could be greatly reduced.
I'm not convinced consciousness is emergent, I don't really have an opinion on _that_- but I'm > 50% convinced that consciousness itself doesn't require a neural network as large as a human brain's.
Transistors are already much smaller than neurons. And of course the brain doesn’t have a clock. And neurons have more complex behavior than single transistors… The whole system is just very different. So, this doesn’t seem like a strategy to get past a boundary, it is more like a suggestion that we give up on the current path and go in a radically different direction. It… isn’t impossible but it seems like a wild change for the field.
If we want something post-cmos, potentially radically more efficient, but still familiar in the sense that it produces digital logic, quantum dot cellular automata with Bennett clocking seems more promising IMO.
1) Noise is an issue as the system gets complex. You can't get away with counting to 1 anymore, all those levels in between matter. 2) Its hard to make an analog computer reconfigurable. 3) Analog computers exist commercially believe it or not, but for niche applications and essentially as coprocessors.
Not sure who’s working on that but I can’t believe it’s not being examined.
https://www.analog.com/en/resources/analog-dialogue/articles...
Our brain has a pretty bounded need of scaling, but once we create some computer equivalent, it would be very counterproductive to make it useless for larger problems for a small gain on smaller ones.
Yes!
> Our brain has a pretty bounded need of scaling
No!
Over aeons our brains scaled from several neurons to 100 billion neurons, each with 1000 synapses. They were able to do it because our brains are digital. They lean on their digital nature even more than computer chips do.
Action potentials are so digital it hurts. They aren't just quantized in level, but in the entire shape of the waveform across several milliseconds. Just as in computer chips, this suppresses perturbations. As long as higher level computation only depends on presence/absence of action potentials and timing, it inherits this robustness and allows scale. Rather than errors accumulating and preventing integration beyond a certain threshold, error resilience scales alongside computation. Every neuron "refreshes the signal," allowing arbitrary integration complexity at any scale, even in the face of messy biology problems along the way. Just like every transistor (or at least logic gate) "refreshes the signal" so that you can stack billions on a chip and quadrillions in sequential computation, even though each transistor is imperfect.
Digital computation is the way. Always has been, always will be.
They also suffer from the global optimisation problem for layout of calculations so compile time is going to be insane.
Their WSE technology is also already obsolete - Tesla's chip does it in a much more logical and cost effective way.
More than that arguably. CUDA cores are more like SIMD lanes than CPU cores like cerebras's usage of 'core'. Since they have 4 wide tensor ops on cerebras, there's arguably 3.6M CUDA equivalent cores.
And, 9 trillion flops per core in 4.4 million transistors per core. That sounds a bit too good to be true.
You would need 64 of these to get 8 exaflops.
https://www.tomshardware.com/tech-industry/artificial-intell...
Your number is off by 64x.
It can do 125 petaflops at FP16
https://www.tomshardware.com/tech-industry/artificial-intell...
From what I could tell from Nvidia's recent presentation, Nvidia works directly with OpenAI to test their next gen hardware. IIRC they had some slides showing the throughput comparisons with Hopper and Blackwell, suggesting they used OpenAI's workload for testing.
H100's have been generally available (not a long waitlist) for only several months, but all the big players had them already 1 year ago.
I agree with you, but I think you might be 1 generation behind.
> OpenAI used H100’s predecessor — NVIDIA A100 GPUs — to train and run ChatGPT, an AI system optimized for dialogue, which has been used by hundreds of millions of people worldwide in record time. OpenAI will be using H100 on its Azure supercomputer to power its continuing AI research.
March 21, 2023 https://nvidianews.nvidia.com/news/nvidia-hopper-gpus-expand...
Interface hardware, being is perceptible to the senses, gets credit over software.
E.g. when people experience a vivid, sharp high resolution display, they attribute all its good properties to the hardware, even if there is some software involved in improving the visuals, like making fonts look better and whatnot.
If a mouse works nicely, people attribute it to the hardware, not the drivers.
If you work in hardware, and crave the appreciation, make something that people look at, hear, or hold in their hands, not something that crunches away in a closet.