That aside, we would need to see some evidence of AI developments being bootstrapped by the previous SOTA model as key part of building the next model.
For now, it's still human researchers pushing the SOTA models forwards.
When people use the term exponential I feel that what they really mean is 'making something so _good_ that it can be used to make the N+1 iteration _more good_ than the last.
https://www.lesswrong.com/posts/qLe4PPginLZxZg5dP/almost-all...
Where's the falsifiable framework that demonstrates your conclusion? Or are we just supposed to trust your intuition?
I like that the HN crowd wants to believe AI is hype (as do I), but it's starting to look like wishful thinking. What is useful to consider is that once we do get AGI, the entirety of society will be upended. Not just programming jobs or other niches, but everything all at once. As such, it's pointless to resist the reality that AGI is a near term possibility.
It would be wise from a fulfillment perspective to make shorter term plans and make sure to get the most out of each day, rather than make 30-40 year plans by sacrificing your daily tranquility. We could be entering a very dark era for humanity, from which there is no escape. There is also a small chance that we could get the tech utopia our billionaire overlords constantly harp on about, but I wouldn't bet on it.
Mr. Musk's exitement knew no bounds. Like, if they are the ones in control of a near AGI computer system we are so screwed.
Unfortunately, we seem to be on this exact trajectory. If open source AGI does not keep up with the billionaires, we risk sliding into an inescapable hellscape.
Dunno about Zuckerberg. Standing still he has somewhat slided into the saner spectrum of tech lords. Nightmare fuel...
"FOSS"-ish LLMs is like. We need those.
[0]: https://ourworldindata.org/grapher/exponential-growth-of-par...
[1]: https://ourworldindata.org/grapher/exponential-growth-of-dat...
[2]: https://epoch.ai/blog/trends-in-training-dataset-sizes
[3]: https://ourworldindata.org/grapher/exponential-growth-of-com...
Obviously predicting the future is hard, and we won't know where this stops till we get there. But I think a degree of skepticism is warranted.
It certainly will be sigmoid-shaped in the end, but the top of the sigmoid could be way beyond human intelligence.
I'm not a fan of this meme that seems to be very popular on HN. Someone with knowledge in EE and drivers can easily acquire enough programming knowledge in the higher layers of programming, at which point they can fill the gaps and understand the entire stack. The only real barrier is that hardware today is largely proprietary, meaning you need to actually work at the company that makes it to have access to the details.
Things can be complex without being intelligent.
> Exponentially smarter AI meets exponentially more difficult wins.
Another is that it doesn't seem like intelligence is the main/only bottleneck to producing better AIs right now. OpenAI seems to think building a $100-500B data center is necessary to stay ahead*, and it seems like most progress thus far has been from scaling compute (not to trivialize architectures and systems optimizations that make that possible). But if GPT-N decides that GPT-N+1 needs another OOM increase in compute, it seems like progress will mostly be limited by how fast increasingly enormous data centers and power plants can be built.
That said, if smart-human-level AGI is reached, I don't think it needs to be exponentially improving to change almost everything. I think AGI is possibly (probably?) in the near-future, also believing that it won't improve exponentially doesn't ease my anxiety about potential bad outcomes.
*Though admittedly DeepSeek _may_ have proven this wrong. Some people seem to think their stated training budget is misleading and/or that they trained on OpenAI outputs (though I'm not sure how this would work for the o models given that they don't provide their thinking trace). I'd be nervous if it was my money going towards Stargate right now.
Babies are born with a fully functioning image recognition stack complete with a segmentation model, facial recognition, gaze estimator, motion tracker and more. Likewise, most of the language model is pre-trained and language acquisition is in large part a pruning process to coalesce unused phonemes, specialize general syntax rules etc. Compare with other animals that lack such a pre-trained model - no matter how much you fine-tune a dog, it's not going to recite Shakespeare. Several other subsystems come online in the first few years with or without training; one example that humans share with other great apes is universal gesture production and recognition models. You can stretch out your arm towards just about any human or chimpanzee on the planet and motion your hand towards your chest and they will understand that you want them to come over. Babies also ship with a highly sophisticated stereophonic audio source segmentation model that can easily isolate speaking voices from background noise. Even when you limit yourself to just I/O related functions, the list goes on from reflexively blinking in response to rapidly approaching objects to complicated balance sensor fusion.
Can we do better than evolution? Probably; evolution is a fairly brute force search approach and we are pretty clever monkeys. After all, we have made multiple orders of magnitude improvements in the state of the art of computations per watt in just a few decades. Can we do MUCH better than evolution at finding efficient intelligences? Maybe, maybe not.
So six billion bits since two bits can represent four values. Base pairs and bases are effectively the same because (from the link) "the identity of one of the bases in the pair determines the other member of the pair."
There was also a startup selling/renting bitcoin miners that doubled as electrical heaters.
The problem is that computers are fundamentally resistors, so at most you can get 100% of the energy back as heat. But a heat pump can give you 2-4 times the energy back. So your AI work (or bitcoin mining) plus the capital outlay of the expensive computers has to be worth the difference.
But if we can make computers that run at, say, 2000 degrees, without using several times more electricity, then we can capture their waste heat and turn a big portion of it back into electricity to re-feed the computers. It doesn't violate thermodynamics, it's just an alternative possibility to make more computers that use less electricity overall (an alternative to directly trying to reduce the energy usage of silicon logic gates) as long as we're still well above Landauer's limit.
Or more on topic see the improvements in LLMs since they were invented. At first each release was an order of magnitude better than the last (see GPT 2 vs 3 vs 4), now they’re getting better but at a much slower rate.
Certainly feels like being at the top of an S curve to me, at least until an entirely new architecture is invented to supersede transformers.
Take the trajectory of chess. handcrafted rules -> policies based on human game statistics -> self-play bootstrapped from human games -> random-initialized self-play.
And if the improvements it makes are not asymptotically diminishing.
If that is a normal human estimation I would guess in reality it is more likely to be in 6-10 years. Which is still good if we get it in 2030 - 2035.
And this same RL is also creating improvements in small model performance.
So, more LLMs are about to rise in quality.
If yes, then you get exponential increases very trivially. If no, then something external continues to bottleneck progress.