A stubborn computer scientist accidentally launched the deep learning boom
arstechnica.com
arstechnica.com
For an extreme example, I recommend Veratasium's video on blue LEDs, https://www.youtube.com/watch?v=AF8d72mA41M. But despite how extreme that example is, it does seem to be a trend. The direction that the herd has gone has been generally well-explored simply because a herd of researchers have already gone there. If they've found nothing, you'll probably find nothing either. So you have to back up.
Sadly, the way that academia is structured discourages this at every turn. :-(
Academia the calling is the pursuit of knowledge, truth, the advancement of humankind. It sometimes needs flexibility, other times rigor. Sometimes a full team of aligned individuals. Sometimes the singular concentration of one person over a long period of time. Results are not guaranteed, and even when they happen they may be revolutionary, or they may be a tiny stepping stone that only reveals it's greater purpose literally CENTURIES later. (Joseph Fourier died in 1830, never realizing that 2 centuries from his all modern life would not be possible without mathematical concepts named after him)
All of this is great, but...how does one survive to be able to do this? Historically, virtually every famous scientist you can think of was a noble. They had family money, servants. They didn't have to think about rent or laundry or cooking or parenting.
If instead we want to be able to guide humans with the ABILITY to do great things towards such pursuits, how do we compensate them? How do we track progress, and create reward structures to incentivize both achievement and progress?
Therein lies academia the profession. I am not in it, nor have I ever been it. I have nothing to gain from defending it. Every ex-academic in industry I come across is so relieved to be out of it.
But that is so sad and disappointing to me, and it pains me how often it's trashed on hackernews...
However I have to say that Sabine Hossenfelder has done an excellent job of showing that one can still follow academia the calling by becoming a youtuber. I just wish that that role had more slots available for the intellectually honest...
The last great scientist who I know to have been a paragon of intellectual honesty was Richard Feynman. Maybe that reflects my ignorance. I think that it reflects how far academia has fallen. With prestigious positions such as President of Stanford University being won by outright fraud.
You can be the judge...
- The ImageNet dataset / competition pre-dates neural net entrants, and I'm not aware that Fei Fei created it in anticipation of such. The reason that AlexNet (2012 ILSVRC/ImageNet entry) made such an impact was that it beat all non-ANN entrants by such a huge margin that it was impossible to ignore, and pretty much killed all ongoing attempts to hand-design transform-invariant image features (such as SIFT).
- While NVidia deserve credit for enabling GPU-compute with CUDA, it was pure luck that ANNs subsequently took off, and became the primary use of CUDA (which then added ANN specific libraries such as cuDNN).
- AlexNet (ImageNet 2012) was certainly what started the ANN boom, which Hinton/LeCun/Bengio then decided to re-brand as "deep learning" to escape any historically negative association of "neural networks". However, while AlexNet demonstrated the power of a large (& deep) neural net, I don't think it's fair to say that it was responsible the current/waning belief in LLM "scaling laws". The immediate aftermath of AlexNet wasn't bigger datasets, but attempts to build bigger/better ANNs to do better on the ImageNet benchmark. The LLM "scaling laws" originated from OpenAI's GPT-1 and GPT-2 where unexpected model capabilities lead to experimentation of scaling, with Sutskever and Amodei being two of the earliest believers. Sutskever has recently said that he thinks that transformer scaling has plateaued, and has started his own company (SSI) pursuing a different approach.
I don't think we can say that Hinton accidentally created the deep learning boom. He always (to his huge credit) believed in ANNs, pushed it into the public eye with AlexNet, created the "Deep Learning" branding, and generally promoted it until it got too powerful for his liking.
> I don't think we can say that Hinton accidentally created the deep learning boom.
Yeah, the headline is unclear, but I read it as referring to Fei Fei being the "accidental" contributor to deep learning, not Hinton.
This is a good example of how investor behaviour can only quantitatively project what will happen in the future. Huang's bet on GPUs for high performance computing made sense in the long term.
Intel didn't have the staying power with the i860[1] a decade earlier (and of course had no idea how to offer decent developer tools). I tried really hard to develop meaningful and executable programmes with an 8bit card (DSM). CUDA was a revelation for me.
But then I realized the author might have been a toddler at that time so might not have lived through it, making it easier to forget.
How time flies...
[1]: https://en.wikipedia.org/wiki/2007%E2%80%932008_financial_cr...
Original discussion: https://news.ycombinator.com/item?id=42057139
https://news.ycombinator.com/item?id=42116140
https://news.ycombinator.com/item?id=42106623
You can check the past discussions of the same article using the past link underneath the title post.
The deep learning boom caught almost everyone by surprise - https://news.ycombinator.com/item?id=42057139 - Nov 2024 (188 comments)
This is just plain not true.
Who edited this? Did no one even bother to skim the Wikipedia page on GPUs?
chatgpt, probably.
Unless I'm getting all of that wrong, it doesn't sound like that bad of a term. I get that it has since been used and abused to absurdity. But it's not like it came out of nowhere at first.
"3-D graphics systems maker nVidia (NVDA) Tuesday unveiled its next generation graphics accelerator, which it hopes will deflect some of the attention from new console gaming machines. The GeForce 256, which nVidia calls the world's first graphics processing unit, will be the successor to its TNT 2 line of graphics chips."
> https://money.cnn.com/1999/08/31/technology/nvidia/
Reading deeper, it seems like Nvidia considered it the first because it was the first GPU to ship with hardware accelerated T&L:
"GeForce 256 was marketed as "the world's first 'GPU', or Graphics Processing Unit", a term Nvidia defined at the time as "a single-chip processor with integrated transform, lighting, triangle setup/clipping, and rendering engines that is capable of processing a minimum of 10 million polygons per second"."
> https://en.wikipedia.org/wiki/GeForce_256
I would also consider earlier NVidia parts like the TNT2 to be GPUs, personally, not sure I agree with this claim either. I certainly called my TNT2 a graphics card back when I bought one in the 90s.
* The first wheel with an axle**.
** The first wheel with an axle for horseless carriages.
"He then reached out to Nvidia. “I sent an e-mail saying, ‘Look, I just told a thousand machine-learning researchers they should go and buy Nvidia cards. Can you send me a free one?’ ” Hinton told me. “They said no.”"
to Altman saying he needs $7T and people eagerly lining up to give the money. Moore law for the venture money in tech.
And each time the new tech wave is higher. Somebody somewhere is already working on that next "trillions dollar something in 10-15 years". Quantum computing based NNs? Or may be just robots everywhere.
In middle school I estimated to rival the processing power of a human brain you would need the processing power of an Empire State building full of Pentium 3 processors. It was a pretty rough calculation. I'm having a hard time remembering just how it went.
lol
"Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should." —"Dr. Ian Malcolm" (Jeff Goldblum), Jurassic Park