I sensed anxiety and frustration at NeurIPS 24
kyunghyuncho.me
kyunghyuncho.me
Why did Brain Exist? https://www.moderndescartes.com/essays/why_brain/ Who pays you? And why? https://www.moderndescartes.com/essays/who_pays_you/
Free lunch briefly existed for a small lucky few in 2017-2021, but today there is definitely no more free lunch.
> once the first generation of lucky PhD’s (including me!) who were there not out of career prospects but mostly out of luck (or unluck), we started to have a series of much more brilliant and purpose-driven PhD’s working on deep learning. because these people were extremely motivated and were selected not by luck but by their merits and zeal, they started to make a much faster and more visible progress. soon afterward, this progress started to show up as actual products.
I have to say, though, that AI and robotics are going through similar transitions as were highlighted in TFA. Robotics has been basically defined by self driving cars for a long time, and we're starting to see the closure of some very big programs. Already the exorbitant salaries are much lower based on what I've seen, and demand is flat lining. My hope is that the senior level engineers with good PhD backgrounds move out into the broader field and bring their experience and research zeal with them as a force multiplier. I expect the diaspora of talent to reinvigorate industry innovation in robotics.
So it will be with LLM-focused researchers in industry in the next phase after we pass peak hype. But the things those battle-scarred researchers will do for the adjacent fields that were not hype-cycled to death will probably be amazing.
Unless they succeed in replacing themselves with general AI. Then all bets are off.
RTFA == "you should have actually read the article instead of wasting everyone's time"
TFA == "referring specifically to the [original, in this context] article"
https://en.m.wikipedia.org/wiki/Elite_overproduction
Since 2012 (Alexnet winning Imagenet with DL) “AI” has been dominated by corporations for one simple and obvious reason that Suttom has pointed out over and over: The group with the best data wins
That wasn’t always true. It used to be the case that the government had better data or academia had better data, but it’s not even close at this point to the point where both of those groups are nearly irrelevant in either applied AI or AI research.
I’ve been applying “AI” at every possible chance I have since 1998 (unreal engine + A* + knn …) and the field has foundationally changed three or four times since.
When I started, people looked at those of us that stated out loud that we wanted to work on AGI as totally insane.
Bengio spoke at the AGI 2014 conference, and at a lunch between myself, Ben G, Josha Bach and Bengio we all agreed DL was going to central to AGI - but it was unclear who would own the data. Someone argued that AGI would most likely come out of an open source project and my argument was there are no structured incentives to allow for the type of organization necessary at the open source level for that.
My position was that DL isn’t sufficient - we need a forward looking RL streaming reward system that is baked into the infrastructure between planning and actuation. I still think that’s true and why AGI (robots doing things instead of people at greater than human level) is probably only a decade away.
It’s pretty wild that stating out loud that your goal is to build superhuman artificial intelligence is still some kind crazy idea - even though it’s so obviously the trajectory we’re on.
As much as I dislike OpenAI and the rest of the corporate AI companies, I respect the fact that they’ve been very vocal about trying to build general intelligence and superhuman intelligence.
I agree, but don't you think some degree of symbolism or abstraction building is necessary for AGI? Current systems seem to be too fragile and data-intensive.
Yes. In 1900, the thing to be was an expert electrician. I had a friend with a PhD in bio who ended up managing a coffee shop because she'd picked the wrong branch of bio.
LLMs may be able to help with productizing LLMs, further reducing the need for people in that area.
(Some physicists told me about how quickly Ken Wilson's application of RG to phase transitions went from the the next big thing to old hat, for instance.)
When I can bear to read editorials in CACM I see the CS profession has long been bothered by whipsawing demand for undergraduate CS degrees. I've never heard about serious employment problems for CS PhDs and maybe I never will because they have a path to industry that saves face better than the paths for physics.
Maybe we will hear about a bust this time. As a cog in the social sciences department, I used to have a view of a baseball diamond out my office window but now there is a construction site for a new building to house the computer science, information science and "statistics and data science" departments which are bulging in undergraduate enrollment.
Will there finally be a bust?
> Deep learning and LLMs advanced so quickly that a large bust in AI didn’t really occur.
I think the problem is we railroad too much. It’s LLMs/large models or bust. It can be hard to even publish if you aren’t using one (e.g. build on a pertained model). The problem is putting all our eggs in one basket when we should diversify. The big tech companies hired a lot of people to freely research but research just narrowed. We ignored the limitations that many discussed for a long time and we could have solved them by now if we just spent a small portion of the time and money we have on railroad topics.What I’ve seen is research become very product focused. Fine for industry research but we have to also have the more academic research. Academic is supposed to do the low TRL (1-4) while industry does higher (like 4-6). But when it’s all mid or high you got nothing in the pipeline. Even if the railroad will get us there there’s always other ways and ways to increase efficiency.
This is a good time to have this conversation as we haven’t finished CVPR reviews. So if you’re a reviewer, an AC, or meta, remember this when evaluating. Evaluate beyond benchmarks and remember the context of the compute power people have. If you require more than an A100 node you only allow industry. Most universities don’t even have a full node without industry support.
https://www.nasa.gov/directorates/somd/space-communications-...
This is different. It’s a weird combination of huge amounts of capital chasing a very specific idea of next token prediction, and a slowdown in SWE hiring that may be related. It’s the first time “AI” as an industry has eaten itself.
I don’t think the hiring slowdown and AI are related. Some companies are using rhetoric about AI to save face but the collapse in the job market was due to an end to ZIRP.
https://www.businessinsider.com/ai-down-rounds-rise-valuatio...
> a lot of these PhD’s hired back then were therefore asked to and free to do research; that is, they chose what they want to work on and they publish what they want to publish. it was just like an academic research position however with 2-5x better compensation as well as external visibility and without teaching duties,
exactly!
> such process standardization is however antithetical to scientific research. we do not need a constant and frequent stream of creative and disruptive innovations but incremental and stable improvements based on standardized processes.
A lot of the early wave AI folks struggle with this. They wanna keep pushing wild research ideas but the industry needs slow incremental stuff, focused on serving
Even being in elite research groups at the most prestigious companies you are evaluated on product and company Impact, which has nothing to do with how groundbreaking your research is, how many awards it gets, or how many cite it. I had colleagues at Google Research bitter that I was getting promoted (doing research addressing product needs - and later publishing it, "systems" papers that are frowned upon by "true" researchers), while with their highly cited theoretical papers they would get a "meet expectations" type of perf eval and never a promotion.
Plus, there were quite a few places where a good publication stream did earn a promotion, without any company/business impact. FAIR, Google Brain, DM. Just not Google Research.
DeepMind didn't have any product impact for God knows how many years, but I bet they did have promos happening:)
And if you join as a PhD fresh grad (RS or SWE), L4 salary is ok, but not amazing compared to costs of living there. From L6 on it starts to be really really good.
People who don't contribute to the bottom line are the first to get a PIP or to be laid off. Effectively the better performers are subsidizing their salary, until the company sooner or later decides to cut dead wood.
You know that in academia you constantly have to beg for money by trying to convince government agencies that you’re bringing them value right?
That was an exaggeration. No employee has full freedom, and I am sure it was expected that you do something which within some period of time, even if not immediately, has prospects for productization; or that when something becomes productizable, you would then divert some of your efforts towards that.
One of many reasons why Google invented Transformers and many components of GPT pre-trainint, but ChatGPT caught them "by surprise" many years later.
I know Feynman was somewhat critical to IAS, and stated that the lack of accountability and commitment could set up researchers to just follow their dreams forever, and eventually end up with some writers block that could take years to resolve.
They very high salaries are central to the situation.
If you remove high salary then you have a lot more freedom. The tradeoff is the entire point of discussion.
I wonder... There are some academics who are really big names in their fields, who publish like crazy in some FAANG. I assume that the company benefits from just having the company's name on their papers at top conferences.
for research that translates directly to an industry setting, look someplace like KDD. venues like that are where data science/ML teams who do some research but are mainly product focused tend to publish.
1. https://www.investopedia.com/terms/b/boom-and-bust-cycle.asp 2. https://www.investopedia.com/bullwhip-effect-definition-5499...
The world doesn't owe you anything, even if you did a PhD in AI.
Sorry if you picked the wrong thing, but it's the same for anything at degree level.
I could simply respond with "sorry to hear you upset yourself by considering these people entitled", and it would have about the same merit.
It's like telling somebody after a close relative or some else dear to them died that what, did they expect that person would live forever? No, do you?
Do you genuinely think their feelings of betrayal stem from an unreasonable notion of the world? Have you at all considered that the expectations they harbored were not 100% of their own creation?
those "affected" have the same little compassion towards many millions of others who are not lucky to work on AI in top school and visiting top AI conference. That's what makes them entitled.
This is just blatantly painting high profile individuals with a negative brush because it personally appeals to your fantasies.
feeling betrayed is my test criteria, since many others are "betrayed" much more, but those individuals focused on their feelings and not on systematic issues in general, hence this makes them "entitled".
> This is just blatantly painting high profile individuals with a negative brush because it personally appeals to your fantasies.
you have to check your language if you want to continue productive discussion.
> you have to check your language if you want to continue productive discussion.
I disagree that my use of language was unreasonable. And to clarify, I do not wish to continue this conversation, productively or otherwise.
bye then
- make a point A
- oh, counterexample with point B
- ehh that's a good point logically...
but I'll just go with point A anyways, especially when point A is something generally optimistic, and point B is cynical.
I don't like going into social arguments for this reason - it's very easy (and often logically correct) to create an ad-absurdum counter-argument to a lot of our general social beliefs.
But to be a functioning human being, sometimes you just have to take the optimistic choice regardless. I know that certainly when I was constantly listening to B, I was depressed out of my mind; would rather be "delusional and stubborn" on some things than depressed.
Twice two might not be five, but it keeps your sanity.
the PhD is not like a master's program, it's an apprenticeship relationship where you essentially work for a senior professor, and your work directly financially benefits the department as well.
so there is a deal here, but if one end of it is collapsing, the students in some departments might very well feel betrayed -- not by society but by their advisors and programs.
Researchers at big tech companies always struggled to get promoted without showing some sort of impact.
What changed is before they could just publish papers (with 0 reproducibility due to “ai ethics”) and still coast by with a decent salary.
ChatGPT ended the whole “the paper is the product” phenomenon, and the end of ZIRP ended the smooth run for slackers in big tech overall.
AI phds are still in an amazing position compared to any other new grad software engineer.
It's all that you learn doing it, it is not just knowledge, it's the scientific method, social skills, learning to communicate science.
A PhD is a personal path, an experience in life that changes it profoundly, I hope that those students will be able to handle any kind of future issues, not just a very niche applicative aspect of deep learning.
I get the same feeling going to big ML conferences now. The incentives are all about job prospects and hype, rather than doing cool stuff.
The sweet spot is the 200 person conferences, imo. You have a much better focus on the science, can literally meet everyone if you put your mind to it. And the spaces tend to be more about the community and the research directions, rather than job prospects.
For example, the amount of compute needed to obtain state of the art on a benchmark is only obtainable by big labs. How can we improve the learning efficiency of deep learning so that one can obtain similar performance with 100x less compute.
This is the first time I've seen it in an actual article, however.
>this post will be more of less a stream of thoughts rather than a well-structured piece
When you enter a PhD program, you are promised nothing else but hard work (hopefully within a great research group). I don’t know what the author is referring to here.
I'd imagine the most significant aspect of this betrayal of PhD students, as far as it exists, is to polarise the domain along research lines that never very clearly mapped to the perennial objectives of modelling or mere curve-fitting. Had a thread of 'modelling science' been retained then it would always be useful.
Businesses may, at the moment, appear to have been confused into the value of curve-fitting --- the generic skills of algorithm design, scientific modelling, and the like, will survive the next wave of business scam.
Phd hires will continue ofcourse, its the volume and selection criteria that get adjusted. When a technical domain gets standardized its natural that the ratio of PhD's to Masters in a team lowers.
Mostly teams with a mandate to innovate (or at least signaling that intention) will have more of the former. A succesful Phd will always be a strong filter for certain types of abilities and dispositions. Just don't expect stampedes to last forever.
Never mind . . .
I respect your perspective and acknowledge I might be missing key context since I didn’t attend the conference. That said, one interpretation is that you’re grappling with the rapid acceleration of innovation—a pace you’ve helped shape. Challenges that once seemed distant now feel imminent, creating a mix of excitement and unease, like the moment a rollercoaster drop starts to terrify.
Funnily enough, he has started publishing academic papers in which he uses proper casing.
and, some of these PhD students and postdocs are my own under my supervision
How can you supervise students and write this way? This doesn't make any sense to me.At the end of the day it's just numbers in a database somewhere.
It's actually kind of surprising how seriously people take it. I can easily imagine someone saying to the inventor of downvotes: "Yeah but what happens if people just... ignore the number? Why would anyone actually give a crap?"
And so where, inevitably, some of the worst.
Merry Christmas to all!
Democratization of AI was a joke.
Basically what you’re suggesting is that we admit to stealing data at scale. It would kill model production, because no one would be sure of the provenance of their training data. In all honesty, we can’t even be sure right now.
It'd be like asking if a law was passed making all of the companies on the S&P 500 co-ops owned equally by all Americans would solve the democratization issue.