The average brown rat may use only 60 kcal per day, but the maximum firing rate of biological neurons is about 100-1000 Hz rather than the A100 clock speed of about 1.5 GHz*, so the silicon gets through the same data set something like 1.5e6-1.5e7 times faster than a rat could.
Scaling up to account for the speed difference, the rat starts looking comparable to a 9e7 - 9e8 kcal/day, or 4.4 to 44 megawatts, computer.
* and the transistors within the A100 are themselves much faster, because clock speed is ~ how long it takes for all chained transistors to flip in the most complex single-clock-cycle operation
Also I'm not totally confident about my comparison because I don't know how wide the data path is, how many different simultaneous inputs a rat or a transformer learns from
Only a small part of that 60kcal is used for learning, and for that same 60 kcal you get an actual physical being that is able to procreate, eat, do things and fend for and maintain itself.
Also you cannot compare neuron firing rates with clockspeed. Afaik each neuron in a ml-model can have code that takes several clock cycles to complete.
Also an neuron in ml is just a weighted value, a biological neuron does much more than that. For example neurons communicate using neuro transmitters as well as using voltage potentials. The actual date rate of biological neurons is therfore much higher and complex.
Basically your analogy is false because your napkin-math basically forgets that the rat is an actual biological rat and not something as neatly defined as a computer chip
The conclusion does not follow from the premise. The observed maximum rate of the inter-neuron communication is important, the mechanism is not.
> Also you cannot compare neuron firing rates with clockspeed. Afaik each neuron in a ml-model can have code that takes several clock cycles to complete.
Depends how you're doing it.
Jupyter notebook? Python in general? Sure.
A100s etc., not so much — those are specialist systems designed for this task:
"""1024 dense FP16/FP32 FMA operations per clock""" - https://images.nvidia.com/aem-dam/en-zz/Solutions/data-cente...
"FMA" meaning "fused multiply-add". It's the unit that matters for synapse-equivalents.
(Even that doesn't mean they're perfect fits: IMO a "perfect fit" would likely be using transistors as analog rather than digital elements, and then you get to run them at the native transistor speed of ~100 GHz or so and don't worry too much about how many bits you need to represent the now-analog weights and biases, but that's one of those things which is easy to say from a comfortable armchair and very hard to turn into silicon).
> Basically your analogy is false because your napkin-math basically forgets that the rat is an actual biological rat and not something as neatly defined as a computer chip
Any of those biological functions that don't correspond to intelligence, make the comparison more extreme in favour of the computer.
This is, after all, a question of their mere intelligence, not how well LLMs (or indeed any AI) do or don't function as von Neumann replicators, which is where things like "procreate, eat, do things and fend for and maintain itself" would actually matter.
Have you thought about stepping back from all of this for a few days and notice that you are wasting your time with these arguments? It doesn't matter how fast you can calculate a dot product or evaluate an activation function if the weights in question do not change.
NNs as of right now are the equivalent of a brain scan. You can simulate how that brain scan would answer a question, but the moment you close the Q and A session, you will have to start from scratch. Making higher resolution brain scans may help you get more precise answers to more questions, but it will never change the questions that it can answer after you have made the brain scan.
Num fecisti?
> It doesn't matter how fast you can calculate a dot product or evaluate an activation function if the weights in question do not change.
That's a deliberate choice, not a fundamental requirement.
Models get frozen in order to become a product someone can put a version number on and ship, not because they must be, as demonstrated both by fine-tuning and by the initial training process — both of which update the weights.
> NNs as of right now are the equivalent of a brain scan.
First: see above.
Second: even if it were, so what? Look at the context I'm replying to, this is about energy efficiency — and applies just fine even when calculated for training the whole thing from scratch.
To put it another way: how long would it take a mouse to read 13 trillion tokens?
The energy cost of silicon vs. biology is lower than people realise, because people read the power consumption without considering that the speed of silicon is much higher: at the lowest level, the speed of silicon computation literally — not metaphorically, really literally — outpaces biological computation by the same magnitude to which jogging outpaces continental drift.
Neurons do so much more than a single math operation. A single cell can act as an intelligent little animal on its own, they are nothing like a neural network "neuron".
And note that all neurons act in parallel, so they are billions times more parallel than GPU's even if the operations would be the same.
> A single cell can act as an intelligent little animal on its own, they are nothing like a neural network "neuron".
Unless any of those things contribute to human intelligence, they do not matter in this context.
Cool, sure. Interesting, yes. But only important to exactly the degree to which any of that makes the human they're in smarter or dumber.
To the extent they're independently intelligent, they're the homunculi in Searle's Chinese Room.
> And note that all neurons act in parallel, so they are billions times more parallel than GPU's even if the operations would be the same.
Order of ten million times faster on a linear basis while still a thousand parallel operations.
SpiNNaker needs 100kWh to simulate one billion neurons. So the rat wins in terms of energy efficiency.
> and they tend to be significantly more efficient
Surely you noticed that this claim is false, just from your own next line saying it needing 100 kW (not "kWh" but I assume that's auto-corrupt) for a mere billion?
Even accounting for how neuron != synapse — one weight is closer to a single synapse; a brown rat has 200e6 neurons and about 450e9 synapses — the stated 100 kW for SpiNNaker is enough to easily drive simpler perceptron-type models of that scale, much faster than "real time".
I don't know if the hardware can be scaled up. That's why I wrote "if we're able to scale them" at the root of this thread.
I'm pretty skeptical of the scaling hypothesis, but I also think there is a huge amount of efficiency improvement runway left to go.
I think it's more likely that the return to further scaling will become net negative at some point, and then the efficiency gains will no longer be focused on doing more with more but rather doing the same amount with less.
But it's definitely an unknown at this point, from my perspective. I may be very wrong about that.
Whether those approaches can scale enough to achieve that is relevant to the question, whether the bottleneck is in hardware or software.
I wager that some scrappy resource constrained startup or research institute will find a way to produce results that are similar to those generated by these ever massive LLM projects only at a fraction of the cost. And I think they’ll do that by pruning the shit out of the model. You don’t need to waste model space on ancient Roman history or the entire canon for the marvel cinematic universe on a model designed to refactor code. You need a model that is fluent in English and “code”.
I think the future will be tightly focused models that can run on inexpensive hardware. And unlike today where only the richest companies on the planet can afford training, anybody with enough inclination will be able to train them. (And you can go on a huge tangent why such a thing is absolutely crucial to a free society)
I dunno. My point is, there is little incentive for these huge companies to “think small”. They have virtually unlimited budgets and so all operate under the idea that more is better. That isn’t gonna be “the answer”… they are all gonna get instantly blindsided by some group who does more with significantly less. These small scrappy models and the institutes and companies behind them will eventually replace the old guard. It’s a tale as old as time.
Six million is a start but this tech won’t truly be democratized until it costs $1000.
Obviously I’m being a little cheeky but my real point is… the idea that this technology is in the control of massive technology companies is dystopian as fuck. Where is the RMS of the LLM space? Who is shouting from every rooftop how dangerous it is to grant so much power and control over information to a handful of massive tech companies, all whom have long histories of caving into various government demands. It’s scary as fuck.
Human intelligence improved dramatically after we improved our ability to extract nutrients from food via cooking
https://www.scientificamerican.com/article/food-for-thought-...
We can put a lot more power flux through an AI than a human body can live through; both because computers can run hot enough to cook us, and because they can be physically distributed in ways that we can't survive.
That doesn't mean there's no constraint, it's just that the extent to which there is a constraint, the constraint is way, way above what humans can consume directly.
Also, electricity is much cheaper than humans. To give a worked example, consider that the UN poverty threshold* is about US$2.15/day in 2022 money, or just under 9¢/hour. My first Google search result for "average cost of electricity in the usa" says "16.54 cents per kWh", which means the UN poverty threshold human lives on a price equivalent ~= just under 542 watts of average American electricity.
The actual power consumption of a human is 2000-2500 kcal/day ~= 96.85-121.1 watts ~= about a fifth of that. In certain narrow domains, AI already makes human labour uneconomic… though fortunately for the ongoing payment of bills, it's currently only that combination of good-and-cheap in narrow domains, not generally.
* I use this standard so nobody suggests outsourcing somewhere cheaper.
I think current LLMs may scale the same way and become very powerful, even if not as energy-efficient as an animal's brain.
In practice, we humans, when we have a technology that is good enough to be generally useful, tend to adopt it as it is. We scale it to fit our needs and perfect it while retaining the original architecture.
This is what happened with cars. Once we had the thermal engine, a battery capable of starting the engine, and tires, the whole industry called it "done" and simply kept this technology despite its shortcomings. The industry invested heavily to scale and mass-produce things that work and people want.