But there's a different way to look at it. Because industry went off and solved the now-"boring" problem of building ever-more-powerful computer processors, the people left behind in academia got to invent machine learning, public-key cryptography, modern coding theory, distributed systems... and so on. What makes academic research valuable is not the freedom of "I get to plan my own day", but the freedom of "I get to work on the weird problems that industry doesn't even realize are problems."
In machine learning we already had back-propagation in the 1970, but had to wait another 40 years for Intel and Nvidia to create ever more powerful processors to tackle useful problems.
So I get how they feel - one of the coolest problems ever instantly went from being something anyone could hope to contribute to, to something almost no one can. You can't match the compute to do the things OpenAI & Google & friends do. And if you happen to stumble on something related that can be explored without access to obscene amounts of capital, guess what, the corporate research teams will notice it, and they can do it better than you, and then they can the idea much further than you ever could.
I'll counter that. In the end we need AI that can do training AND inference on edge devices out in the real world. A good (and possibly profitable) example would be robotic pets that can learn (even to understand words) and interact with their owners like real animals, but don't need to go to the vet or eat and poop. Big companies relying on huge compute resources are not even aiming at this type of thing. They're too busy using their "scale" to even bother looking at smaller but interesting methods or solutions.
That's of course if, by this point, we aren't in a middle of a futile scramble to avoid getting extincted by ChatGPT-7 that someone left in self-play mode and forgot to turn off before going on vacation.
Point being, general AI is general. Even at extreme expenditure of resources, the closer the corporations get to it, the more problems they can put it to - including, eventually, the problem of optimizing itself. Already today people are using current-gen models to assist in developing next-gen models; this trend will only continue, until at some point you'll be able to let the model self-improve, mostly unsupervised. I imagine the compute costs per AI value delivered will drop like a stone then.
From what I can tell, the team at OpenAI (following after Google/Deepmind etc.) are simply mashing the pedal to the floor to get bigger and better models from their existing techniques, and then tuning the resulting black box to make it produce more "useful" answers. And that's fine! That's precisely what an industry lab is expected to do: they have the resources to do the training and the need to get products in front of paying customers as quickly as possible to justify it. And frankly with top AI engineers getting paid millions of dollars and Google/Meta tight behind you, emphasizing results is the most viable strategy. If "turn the needle on the box to the right" is giving you good answers, why would you waste a $1-$5m-salary engineer on academic questions like "why does the box do that?"
And yet, asking questions like "why does the box do that?" is the reason technology didn't stop at the steam engine. I suspect that finding the answer to those questions won't immediately sell enterprise licenses, but will be very important. And the answers will probably fall to someone who's making $30k/year in a graduate program.
ETA: Of course, it may turn out that "turn the needle on the box to the right" is enough to obsolete all human researchers, in which case I'll be wrong about this. But it'll hardly matter in that case ;)
I'm only an amateur in the field, so my uneducated high-level understanding is that, in a sufficiently high-dimensional latent space, there's more than enough dimensions to assign to any single semantic relationship people ever thought of, which is what the training process effectively does, which reduces an important part of thinking - working with concepts and their relationships - entirely to vector adjacency search.
I'm only beginning to study the details, and I don't know how much of specific understanding of this exists, but at the very least this high-level model explains why scaling makes qualitative difference here.
Now, I agree they have strong commercial incentives to push their models as far as possible as fast as possible, but honestly, if I were a researcher working on these models, even if I was somehow unconcerned with any kind of commercial viability and had access to more compute, I'd absolutely keep scaling those models up and up, all the way until I hit the limit of available compute, or the models stop qualitatively improving with scale.
Basically, there's no reason[0] to stop now and try to fully comprehend how GPT-2 works, when GPT-3 was a qualitative jump, and GPT-4 even more so, and GPT-5 is around the corner, and GPT-6 might be a year away from now. All those steps yield important new insights into how the whole architecture works, and if at some point the scaling breaks, that would be even more important knowledge to have. And this doesn't even take into account the fact that, starting with GPT-3, those models are increasingly useful in accelerating both research and scaling alike.
----
[0] - Except, of course, that if transformer models are the road to generic AI, then we'll just blindly race straight into a point of no return.
Not to mention that there is zero appetite from undergrads or postgrads to get into the nitty-gritty of it. To learn CNNs at the deep-dive level you need calculus, at least differentiation and integration. Calculus or even pre-calculus doesn't form part of the degree programme for most compsci BScs any more, because it is 'too hard'.
The way most students 'learn' AI is to use a method out of a Python library with near-zero understanding of how it works, and regurgitate it for an assessment.
Professorial research staff in most UK universities are light-years from AI within industry, and there's no clear path to that gap tightening, especially while universities are being run like second-rate consulting houses (don't get me started on THAT).
anyway i think your statement that industry is light years away from unis is just misleading. i think the two are trying to answer different questions: 1. how can i achieve a "somewhat" decent chatbot that gets me rich albeit not even knowing what it does [industry in case you wondered] 2. try to understand, quantify and measure how well a model works, is it stable? does it converge if we have small datasets? and so on so forth.
just my two cents, to conclude i think a good analogy to the current climate is the 700-800s with electromagnetism: plenty of people discovered "empirical" laws but didn't understand really the phenomenon.
Sounds dead on. Do these large """language""" models actually even implement any concepts from linguistics? Or is the entire "language" part of the model merely derived from the fact that it's inherently part of the training data?
I don't fault Chomsky at all for being fed up with the hype here.
The entire field is also glossing over the fact that other languages which aren't English exist.
GP here is, IMO, confusing what the corporations want (1), with what corporate R&D people want (2). As long as the corps see good ROI on throwing infinite money at their AI R&D departments, then those corporate researchers are better positioned and better equipped to do actual, solid science, than academia ever can be. This has happened many times before, including in this industry. Research is best done by well-funded teams of smart people left to do whatever they fancy. When those conditions arise, progress happens, and it doesn't matter whether it's the government or industry that creates them.
(Conversely, the best hope for academia to become relevant again is that corporations lose interest in this research, and defund their departments. This could happen if e.g. transformers end up being a dead end, or compute suddenly becomes very expensive.)
> Do these large """language""" models actually even implement any concepts from linguistics? Or is the entire "language" part of the model merely derived from the fact that it's inherently part of the training data?
The latter. And guess what, they're not trying to solve the issue of linguistics. They started as tools to generate human-sounding text, but in the process of just throwing more data and compute at them, they not only got better, but started to acquire something resembling concept-level understanding.
It turns out that surprisingly many aspects of thinking seem to reduce well to proximity search in a vector space, if that space is high-dimensional enough. This result is both surprising and impactful well beyond the field of AI. It's arguably the first potential path we identified that the evolution could take to gradually random-walk itself from amoeabas to human brains.
btw, in other languages i guess it is decent although it depends on which language, at least gpt4.
What do you mean? Processing a CNN layer takes an amount of time that does not depend on the input data, only the input/output sizes. Fourier transform is just a change of basis. Why should anything speed up?
This seems to be done in some cases. I guess it isn't done more widely because the "standard" convolution kernels are very small and the performance would actually be worse?
I think your complexity argument is correct for N=pixels=kernel size. But typically, pixels>>kernel size.
Disclosure: I work at Arm optimising open source ML frameworks. Opinions are my own.
In the US I’ve never seen a BS in computer science that didn’t require calculus. I can’t speak for the UK, but it would surprise me that what you say is true.
Oh the poor dears, imagine needing schoolboy maths to do science.
That's interesting, could you please elaborate?
You can easily see a lot of people who work at OpenAI on LinkedIn. University AI labs are always left out because non-commercial products just aren’t as noticeable.
I highly doubt that many of these academics would struggle being amount candidates at DeepMind, OpenAI and FAIR.
Prestige is very subjective, which makes it also very broad thankfully.