The AI research job market
interconnects.ai
interconnects.ai
MLOps leads/lags research depending on your application patterns so it’s an extremely dynamic place to be to see what’s happening
I’d argue based on what I’m seeing with implementations, and importantly how FLEXIBLE transformers seem to be, this is the most true part of this article:
“we’re going to get way further with the Transformer architecture than most ideas in the past”
How did you get into this? Seems like a lot of places are stuck on the idea if you didn't do it in the past you cant do it now
Take some courses and get some certifications. And also make some serious projects where you demonstrate your capabilities with cutting edge tools.
This is more focused on tools and use of said tools.
Take some trained models, and demonstrate how well you can use them.
Some ideas:
1. Take a cats vs. dogs model, deploy it online. Design an API around it. Document the API well. Create a mechanism to show confidence score, and store low confidence score examples in a database that you can later manually label and retrain the model with.
2. Take a smallish LLM, design a VS code extension that documents your functions based on docstring.
Just demonstrate your basic knowledge in ML, and really good software engineering skills, learn the vocabulary well, and then start applying for jobs. It's much better if you have a CS/EE degree.
Certifications will do nothing for you. The harsh reality is only real world experience doing this stuff at scale will help you understand all the complexity involved. There are tons of people trying to hop onto this train after taking a few online courses and it's making it hard to filter down candidate pools.
I do think they can be valuable if they help you learn the basics and get started on a bigger personal project, but not as something to put on your resume.
The problem is a lot of tutorials just show you how to make a Flask/Gradio website (maybe FastAPI) and call it a day. A lot of the experience here is the sort of in the trenches practical stuff that you can't cover in a MOOC (and it's expensive to experiment with GPU clusters). I suspect there are better non-ML courses people could take though.
A sibling commenter mentions that certifications will do nothing for you. They're not exactly wrong, because what ultimately matters is that you can demonstrate your skills. Certifications and to a large extent even degrees mean very little; what matters is that you convince them you know how to do stuff. The best way to convince people you know how to do stuff is to be able to show a list of cool things you actually did. These courses and their certifications may not mean much on their own, but in the course of completing the courses you will develop skills and capabilities you can demonstrate and talk about in your resume and cover letter.
For me, my work in software and AI specifically predates 2012 - blood sweat and tears of going from non-big data statistical forecasting programs (Bayes nets) to big data forecasting (R, Python stat packages) to geometric vision (SURF, HOG etc) to big data CNN & MDP image processing for CNNs (tensorflow) etc…
Like I said, blood sweat and tears
The trick for breaking into something like this is to produce a portfolio of one or more projects you did where you demonstrate experience with it. This means actually doing it yourself, have a repo with notebooks and text explaining how everything works.
AI is definitely not the first field that is like this, where it at first appears only the people already doing it are qualified to do it. I have had to do this quite a few times over the last 30 years to stay relevant. It takes a lot of work to do this, but it's easier than ever to do today. Today the tools you need to break into almost any technical field can be freely downloaded. A couple of decades ago if you wanted to create a portfolio for something the tools were not freely available. For example, vxWorks for embedded systems programming, or Oracle for demonstrating you can administer large databases, or 3D Studio Max or Maya for 3d modelling: all of these tools were expensive enough to be inaccessible to an individual.
But today, you can go do independent work, take courses and get certifications, and create your own body of work that demonstrates you have an understanding of the field.
If you want to start making your own body of work in the field of AI, I suggest starting with these resources:
1. FastAI
2. Deeplearning.ai. Get a certification and put it on your resume.
3. Karpathy Zero to Hero
4. Re-create the technique in the ReAct paper (Reasoning and Acting).
If you want to demonstrate capability in AI in general, proceed in sequence 1-4. If you want to demonstrate capability with LLM's in particular, proceed in sequence from from 4-1.
It's statistics. If you can cast your data to a common format (relatively trivial in cases where they are stored as binary representing numerical values) you can learn from patterns. It is not surprising.
I haven’t been keeping up with the LLM papers because it’s not really my academic interest but I do find them impressive, so maybe this has been figured out but there are two reasons they could accommodate new data modalities really easily: either the sequential data we generate in the real world is more similar than we would have guessed or the hard part isn’t in handling the domain-specific data, but learning to process and predict future signals really well. The former case would be more surprising to me but it is certainly a possibility — most domains have a “language” of sorts, e.g., visual motifs or licks that get passed between musicians. The network could be picking up on those “linguistic” features born out in data it is fed and just needs to alter its vocabulary from words to pixels or whatever. The second case would be my guess. If you have an algorithm that is good at predicting the future based on the recent past, the hard part is done. The rest is just optimizing it for the task (language, sound, video) at hand.
It's extremely impressive that the same algorithm can handle many different types of data with no changes.
If they're ints you can trivially make them floating-point using routines that have existed since those data formats were invented. If they're discrete values you can encode them in ways that make them legible to the models. One of the more recent interesting developments was token embeddings, admittedly, but again this is just an example of taking slightly more abstract representations and turning them into bits representing floating-point values, which has been an established paradigm known as one of the pillars of "feature engineering" since the beginning of the ML field. One-hot encoding is just a special case of token embeddings.
It's amazing that humans figured out how to store data in useful ways, not that models can "figure out" what to do with things they are already capable of ingesting and processing.
Have been interested in this stuff for years. Did my CS project with NN just before everyone started using GPU's ,and a short DS course more recently. Seeing all the marketing people move into space with their prompt cheat sheets on LinkedIn while many tech people are ,ironically, locked out by blackbox recruitment algorithms is maddening.(This particular problem goes far beyond tech jobs though).
Some also seem to be mixing up DS and DE roles a bit, one of the few times I got an interview I had end it and apologize as what they were looking for was a data engineer.
Another was listed as a machine learning role ,when I got the offer it was travelling tech support and paid less. With the promise of undefined ML work later.
Some companies are just tacking irrelevant ML and AI stuff onto job descriptions.
Also so many live coding tests , and that one weird recruiter asking about "skeletons in closets"
100% happened to me once. Wasted hours of my time.
> Some companies are just tacking irrelevant ML and AI stuff onto job descriptions.
Some of them do this deliberately. I have seen this practice in companies targeting junior roles and fresh out of college grads. They hire them with shit pay and promise them ML experience, and then make them do non ML stuff.
The second bait and switch example though went on a lot longer. I had an off feeling about it from the first call.
One guy on call was stifling a laugh the whole time.
They made sure to emphasize they we're offering me a lot of experience, doing me favor essentially.
When they gave me the offer they also requested I send them over a professional photograph of myself. Maybe that's normal in some countries but to me it was the red flag that finally made me notice all the other red flags.
Even if for some bizarre reason we’ve already tapped the maximum potential of transformer architectures and all of this money goes nowhere, compared to all the other ways that society wastes money, I would be fine with calling this a big bet for humanity that didn’t pay off. It doesn’t mean that it wasn’t worth the attempt though.
GenAI is a remarkably useful tool, but its not one step away from an AGI.
I guess the challenge is more to agree on a fitness function to measure the "AGI"-progress" against, but thats a different topic. But in general scaling up the current GenAI tech and parallelize/specialize the models in a multi-generational way _should_ be a safe ticket to AGI, but the time scale is inknown of course (since we can't even agree on the goal definition)
The current LLM's get stuck in loops when a problem is too hard for it. They just keep doing the wrong thing over and over. It's not obvious this sort of ai can "build novel attempts" at hard problems.
I think you'll find that AGI cynics do not agree at all that "engineering a 10x/100x version" of what we have and making it attempt "AGI algorithms 24/7 in an evolutionary setting" is a "safe ticket" to AGI.
The value of potential bank scams that are otherwise illegal was enormous to investors though. Lots of people got extremely wealthy thanks to crypto scams. Then when the legal holes were covered crypto was forgotten extremely quickly since the hype was mostly kept alive by scams.
AI doesn't have nearly as lucrative scams, so I doubt you will see the same investor frenzy.
Maybe you are right about the “frenzy”, but quantitatively speaking the market cap of Big Tech (including and especially nVidia) is probably larger than the crypto scams ever will be.
As a comparison the market cap of crypto is apparently less than the cap of nVidia. Share prices of other tech companies like Microsoft are also inflated due to expectations of AI related returns.
The “frenzy” may not be as insane but the money is definitely there. Especially with the crypto bubbles bursting the money has to go somewhere, unless you honestly believe they ended up in US treasury bonds or sth
That was also true, in AI, of the expert-system hype cycle. And the actual value unlocked was extraordinary, just not at the scale people saw as the potential.
Actually, it was seen as true of all of the hype cycles during the hype cycle, that's what makes it a hype cycle.
> (I mean, what was the theoretical upper limit on the benefit of cryptocurrency for the world? Probably not that much.)
If you believed the people that were as breathless about it as you are about the current AI hype cycle, basically infinite, unlocking ways human potential and interactions, economic and otherwise, are held back by centralized and/or authoritarian systems.
That's what made it a hype cycle.
> It’s quite possible that we already have sufficient computational power and the necessary data for AGI—all we need are the right algorithms.
Yeah, but that's always been true. If software-only AGI is possible, we've always had the data in the natural world, and with no strong theoretical model for the necessary computational power, its always been possible we had enough. What we clearly lacked were the right algorithms (oh, and any reason to believe software-only AGI was possible.)
In the far far future, if we did crack AGI, it's not impossible to believe that specialized hardware modules would be built to enable AGI to interface with a "normal" home computer, much like we already add modules to our computers for specialized applications. Would this still count as software-only AI to you?
I've held for a long time that sensory input and real-world agency might be necessary to grow intelligence, so maybe you mean something like that, but even then that's something not incredibly outside the realm of what regular computers could do with some expansion.
This spider could be evidence of "software based intelligence" in biological brains - it exhibits much more complex behaviors than other animals it's size, more comparable to cats and dogs.
What I mean is that some believe that their brain is "emulating" all parts of the larger "brain", but one at a time, and passing the "data" that comes out of one into the next.
Just a cool thing.
According to the crypto-faithful at the time: solving territorial disputes (Gaza Strip? blockchain solves this!), identity management, bank transfers, payments over the internet with no transaction fees, "the supply chain" (whatever that means), etc. Not as interesting to a layperson as AGI, but if all those (or ANY of those) ended up panning out, crypto would have been a multi-trillion dollar industry and fundamentally transformed vast swathes of modern society.
I do think LLMs are far more useful than blockchain, but claiming "the potential value of this one is extraordinary" is exactly what people said in previous hype cycles.
what metric would you like to use, specifically? double check that its a metric that matches other industries
the market cap of the digital spot commodities? the marketcap of the businesses that use the digital spot commodities? the revenue of all participants and service providers? the volume of all shares and futures and spot trades when sliced down to a submetric that represents 'real' trades? all of the above?
> and fundamentally transformed vast swathes of modern society
thats ...a... goal post. I'm not sure if that's a goal post I would have, its market microstructure plumbing. At best, it modifies capital formation, letting different ventures get funding, which it already has.
and then, what time frame? its a pretty good S-curve from 2009. there is a pretty clear chronology of what delays what, everything that has resulted in a seasonal bubble in crypto comes from a software proposal being ratified that allows it to touch another industry that it previously didn't. Many overlapping similarities to IETF proposals for WWW, but I understand this level of discussion might not reach your circles, the point stands that there are plenty of people in the tech space that had the exact same observation and you and chose to contribute to the proposals that make crypto now more accessible to the next group.
There are plenty of proposals now in many different crypto communities, even ones to make ratification more egalitarian and collaborative.
some turn out to be hits for adoption.
I think it is interesting for people to then use that reality to say crypto hasnt fulfilled any lofty idea they overheard an enthusiast say, because it took too long.
Prior proposals and their ratification were necessary for the reported market cap to reach $1bn, but I know I know “market cap!? you cant sell it all at once!” Holding crypto assets and industry to a separate higher standard than all other industries on the planet.
The same people saying this are often the same ones betting on sovereign currencies crashing.
That sounds exactly like most hype cycles, it's almost a tautology that the perceived potential value is immense (at least to enough people).
Consider e.g. the hype around "the internet" in early mid nineties, which led to the dot.com collapse. Today the internet has undeniably had a massive impact globally, so the naysayers have been comprehensively proven wrong. On the other hand, the most optimistic views have not begun to come to pass yet (ever?) either. Lots of ideas that were floated in the 90s didn't really work until 10, 15, 20 years later. Some things that are now ubiquitous weren't really conceived of then, etc. etc. As usual, it turned out the technology wasn't the really hard part.
So far the current AI cycle seems to be following the usual playbook.
I’d be happy just being a cog in the machine, work 9-5, and get to have an upper middle class lifestyle with my family the rest of the time. That’s probs better than what 95% of people (in the US) get to experience.
There was hype with the Internet, and lots of scams, and naysayer, and bogus money, and real money.
I remember reading a Java 1.0 Book, and someone just casually saying "why learn that, that internet thing isn't going to last, it's all hype."
And how is that different than "AI" ? It's not like these techniques sprang out of the ether in the 2000s.
I read a ton of AI research papers written in the 80s and 90s.
Lack of AI hype is was because the lack of data and compute. The actual field is here for a very long time.
But if we're going to analogize to the internet, you have to count LLMs as "the internet" and the things that came before as just "networking". Networking is very old and laid the basis for the internet, but the internet was something different in kind.
My only point is that LLM being used for real things is very recent. There hasn't been a ton of use (relatively speaking) before it was available for public use. That's not true with the internet. The internet was heavily used, and heavily iterated on, before the public had access to it.
Being skeptical about LLMs is not irrational, and so you see a lot of it. There's just not a great deal of history with them to provide counterexamples to that skepticism. That's not the case when the internet was opened to the public, which is a large factor in why the skepticism about the internet was much lower than the skepticism about LLMs.
Yes, there was usenet, and gopher, and FTP.
Think splitting hairs a bit.
As others have pointed out. 'AI', has been actively researched since the 40's. So there was a lot of ground work before the latest shiny thing the 'LLM'.
Just as there was a lot of ground work in 'networking' before WWW.
I've been talking to GPT4 about a problem all day. I think everyone saying it is 'over-hyped' have really just not used it yet.
It will definitely keep growing, it will replace jobs, it is already increasing productivity and changing markets.
Sorry - one edit. I strongly disagree that things were 'well proven' with the Internet before the public was aware. If you include WWW and the rapid changes in standards and browsers, and technology. It was all moving as fast and with all the bugs and problems and hacks as needed for fast moving tech. There was a ton not proven, and with no use cases to justify it. It was a wild west. Now we are seeing it again.
Sometimes hype becomes reality many years later. This is probably going to be true for AI as well.
I post this to say there is nothing wrong with you. You hear of the crazy successful few but not the majority of cases.
Except engineers don't have a union nor contracts like athletes, so:
* When layoffs / mass firings happen, engineers don't get guaranteed money
* The comp of the top 1% doesn't pull up the bottom 99% (not nearly as fast). Much of the bump in SWE salaries in the past 5 years came from the uncovering of the Apple+other co's no-poaches. The $1-5m retainers Google was offering in 2010-2012 to keep G from jumping to FB look now like peanuts.
* Engineers at competing companies have no say over each other's comp. It takes actual offers for salaries to rise instead of Eng being able to pool the salary data and stats and determine who's Big Head and who all are the 10,000 underpaid SWEs.
Moreover, the glamour of the "transfers" is pretty tightly contained within the author's bubble. A lot of people in AI-adjacent fields still don't even know what Pytorch / Tensorflow are.
Important historical context the author leaves out: Hinton accepting $35-40m for his lab to do deep learning at Google is what set most of the initial benchmark. Many stories out there where Hinton broke his NDA, here is one for example: https://www.gq-magazine.co.uk/culture/article/cade-metz-geni...
It's important context because there's so much arbitrage of private info in AI, you can't take posts like the author's at face value.
If things appear to change that fast, I’d suggest one isn’t well calibrated to the overall environment.
There are deeper fundamentals that don’t change very often. The specific industry moves and product offerings can only deviate so much from ground truth. The more they deviate, the more one should be skeptical of their claims.
Fundamental research tends to be done in academia. Big tech does some of it, but right now they're more focused on making LLMs into products.
A mid-stage Computer Vision software startup just hired its first ML Engineer to work on multi-modal LLM, VLLM, GenAI for image & language-based tasks.
Their product is SLAM/Perception-focused and they have many CV Engineers, yet even they've found a need for LLM.
Does that mean that DM now does no fundamental research - or does it still happen and it has simply been rebranded/hidden away?
Would the bubble have burst before one can finish a PhD?
If this is about your financial outcome, make sure to factor in the opportunity cost of a PhD. It will require 5-7yrs where you will make very little money.
I haven't seen anything indicating it's essential, and I still see that with a MS in AI/ML you're still most likely going to be doing Software Dev... I'm sure it's different for a PhD, but as the other commenter said, it's going to take a lot longer.
So the next hot thing is likely an ML field that is not the current hottest.
It's going to burst for sure, but people will be using ML/AI anyway.
If that's your concern, then you can safely do ML/AI.
But, don't hold your breath on getting a presrigious role after PhD.
Because, there are very few real AI companies. And their hiring is skewed to Stanford, MIT, UCLA-B, Oxbridge, UToronto, ETH, etc. And getting into these schools in AI was always competitive, but now it is crazy-like because all the prep school kids are eyeing this since basically high-school or even before.
So, there is someone like me, who didn't know about modern AI until last year of college in a different major, then learning on my own and getting research jobs in small time companies with shit pay, and then there are people with tiger parents who are white collar or even academics who helps their children get into a really prestigious school, and there they do AI projects with top professors, who write them recommendations, and who also do summer internships in DeepMind, and then they use that to get a job in a proper ML company. At this point they have 2-4 publications in top tier conference. Then they work in BigTech AI lab for 3-5 years and get at least 4/5 more papers (no upper limit). And these are the people who are going to Stanford PhDs. Not people like me. And then these PhDs will be Research Engineers and such in DeepMind, OpenAI, etc.
So, before deciding to do a PhD from non elite institutions, think hard.
Because of AI hype, the situation is very bad for the rest of us. Because all of the prep school types in STEM want to make it in AI.
Research Scientists. You don't need all that pedigree to work as a Research Engineer.
When they talk about sending a video they’re referring to TCP/IP. I bet they can’t even talk about elliptic curve cryptography!
GP here is out of their depth.
You don't need graph theory to understand NNs at all.
You need Linear Algebra and Differential Calculus, though.
You do need graph theory to understand and do Graph Neural Networks. But that's a subfield of modern AI and many AI researchers don't know/study/research Graph NNs at all.
You don't actually need to know graph theory to do RL, Vision, ANNs, NLP, etc. unless explicitly needed in your research/job.
I did some hiring for a very real machine learning (AI if you want to call it that) initiative that started even before the LLM explosion. The number of candidates applying with claimed ML/AI experience who haven’t done anything more than follow online tutorials is wild. This was at a company that had a good reputation among tech people and paid above average, so we got a lot of candidates hoping to talk their way into ML jobs after completing some basic courses online.
The weirdest trend was all of the people who had done large AI projects on things that didn’t need AI at all. We had people bragging about spending a year or more trying to get an AI model to do simple tasks that were easily solved deterministically with simple math, for example. There was a lot of AI-ification for the sake of using AI.
It feels similar to when everyone with a Raspberry Pi started claiming embedded expertise or when people who worked with analytics started branding themselves as Big Data experts.
From physics I have a good theoretical grounding in how ML works (optimizing a cost function over a high dimensional manifold to reconstruct a probability distribution, then using the distribution for some task) but I personally find actually ‘doing ML’ to be rather dull.
I can relate to this a lot. In my company many things you can sell as "AI" can really be solved with traditional data processing.
Fad-chasing often leads to silly technical decisions. Same thing happened with blockchains when they were at the peak of the famous hype cycle. [0]
This is how people get experience with ML though. I don’t think that’s a bad thing.
It sounds like you’re looking for a candidate with current ML experience. But I’ve seen so many people go from zero knowledge to capable devs that this seems like a mistake. You’ll end up overpaying.
Just try to find someone with a burning ambition to learn. That seems like the key to get someone capable in the long run. If they point out something beyond Kaggle that makes you think, pay attention to that feeling — it means they’re in it for more than the money.
I wish there was an easier way to label roles differently based on when you just need to throw X or Y model at some chunk of data and when more specialized modeling is required. Previously it was roughly delineated by "data science" vs "ML" roles but the recent AI thing has really messed with this.
If you're teaching them, you shouldn't be paying them at the AI expert rate.
These days people can get an excellent introductory class to spark and be just as good as I've ever been at it. I wouldn't call them 'charlatans' like the poster above did. It's just that the libraries used to implement spark have been abstracted and people learn it faster.
That's just how it goes in tech. Anyone who wants to learn is a treated like a poser. We over-index on academic credentials which are really not indicative of actual hands-on engineering ability.
PS. There are no AI/ML experts. There are LLM experts, prediction model experts, regression experts, image recognition experts.... If you are hiring a 'AI/ML expert', you have no idea what you are hiring.
But that doesn’t mean that having people with actual research/depthful expertise aren’t essential and hard to find amongst the noise.
The person you responded to is talking about would-be technicians applying for researcher roles. That happens in tech booms and opens amazing doors for lucky smart people, but it’s also a huge PITA for hiring managers to deal with.
One important aspect of all skills is to know the limitations and boundaries of said skill. It’s probably fine if somebody implemented ML on a trivial problem to learn and practice, but if they didn’t realize there could be better solutions in the first place and that ML isn’t a solution to everything, then it’s a big red flag for me.
Also, finding a good problem for a solution is also a handy skill, if one can’t figure out how to apply their skills to a real problem, then that does give slightly negative impressions.
I think the hype on the field and the shitty candidate pool go hand in hand. The shitty candidate pool will groupthink / cargo cult the space without much critical thinking. The groupthink / hype will cause people to jump into the field who don't have any business being in the field.
I've seen two variants of this
1) People that have worked for traditional (as in non-tech) companies, where there's been a huge push for digitalization and "AI". These things come from the very top, and you don't really have much say. I've been there myself.
The upper echelon wants "AI" so that they can tick off boxes to the board of directors. With these folks, its all about managing expectations - but frankly, they don't care if you implement a simple regression model, or spend a fortune on overkill models. The most important part is that you've brought "AI" to the company.
2) The people that want to pad their resumes. There's no need, no push, but no-one is stopping you. You can add "designed and implemented AI products to the business operation blablabla" to your CV.
These days, I've seen and experienced 1) an awful lot. It's all about keeping up with the joneses.
It should be possible for a competent software engineer to get up to speed in AI in less than 6 months and much of that time can be on the job itself.
I don't think AI is hard to learn. The fundamentals are extremely simple and a competent software engineer can learn all the required concepts in a few months. It's easier if you already have a background in mathematics but not required. If you can write software then you can learn how to write differentiable tensor programs with any of the AI frameworks.
Edit: You asked what it is about these jobs that requires expertise. I answered: it requires expertise to create competitive models. So companies that need competitive models requires expertise.
Edit: Why do you ask? I don't see why it is relevant for the discussion.
AI is ill-defined so the premise of your comment makes it difficult to answer. For small well-known tasks (image classification, object detection, sentiment detection) that is train-once on a single dataset and deploy-once what you are saying is true, but for more complex products there is a lot of arcane knowledge that can go in training/deploying/maintaining a model.
On the training side, you need to be able to define the correct metrics, identify bottlenecks in your dataloader, scale to multiple nodes (which is itself a sub-field because distributing a model is not simple) and run evaluation. Throughout the whole thing you have to implement proper dataset versioning (otherwise your evaluation results won't be comparable) and store it in a way that has enough throughput to not bottleneck your training without bankrupting the company (images and videos are not small).
Finally you have a trained model that needs to be deployed, GPU time is expensive so you need to know about compilation techniques/operator fusing, quantization and you need to be able to scale. The requirements to do that are complex because the input data is not always just text.
So yes all the above (and a lot more) require specific expertise.
I know cause I did it.
And I knew the math beforehand. I was a Physics major in college with a CS minor.
Knowing the library is the least hard part about ML work just like knowing the web framework is the least hard part about webdev (both imo). It's much more important to understand the actual problem domain and data and get a smooth data pipeline up and running.
Scaling, optimizing inference, squeezing out better performance and annoying labeling. There's a pretty solid gap from applying some framework to a preexisting and never changing dataset vs. curating said dataset in a changing environment. And if we're talking about RL and not just supervised/unsupervised then building a suitable training environment etc. also become quite interesting.
If someone asked me "what's so hard about webdev" my answer would be similar btw...it's fairly easy to set up a reasonably complicated "hello world" project in any given framework but it gets a lot harder when real world issues like different auth worklflows, security, scaling and handling database migrations etc. enter the picture.
Several years ago on HN there was a blog post which (attempted to) answer this question in detail, and I have been unsuccessfully trying to find it for a long time. The extra facts I can remember about it are:
* It was by a fairly well known academic or industry researcher
* It had reddish graphics showing slices of the problem domain stacking up like slices of bread
* It was on HN, either as a submission or in the comments, between 2016 and 2018.
If anybody knows the URL to this post, I would be stoked!
--
I decided to have a crack myself and see what come back.
Here's a few names / blogs that might be useful:
Chris Olah: He's written extensively about deep learning and AI. His blog, colah.github.io, has a unique graphical style that helps explain complex topics.
Distill.pub: This online journal publishes clear and visually engaging articles on machine learning topics. Some of the articles have been discussed on HN.
Andrej Karpathy: Director of AI at Tesla and previously a researcher at OpenAI and Stanford. He's known for his blog, karpathy.github.io, where he delves into various AI topics.
Ian Goodfellow: Known for inventing Generative Adversarial Networks (GANs) and for his deep learning textbook. He might have some writings that match your description.
Ben Recht: A professor at Berkeley who writes about the challenges and misunderstandings in machine learning on his blog, www.argmin.net.
Sebastian Ruder: He has written many articles about NLP and machine learning at ruder.io.
You can also try searching via https://hn.algolia.com/
But becoming an X-Scientist (Data/Applied/Applied Research) is a whole different skill set. Now, this kind of role only exists in a proper ML company. But, just acquiring the Statistics & Linear Algebra 201 level intuition is about 6 months of fulltime study in its own right. You also need to have deep skills in one of the Tabular/Vision/NLP/Robotics areas and get hired into a role accordingly. Usually 1 year intensive masters level is good enough to get your foot in the door, with the more prestigious roles needing about 2 years of intensive work with some track record of State-of-the-art results on 1 occasion.
Then you have proper researchers, and that might be the most impossible to get in field right now. I know kids who have only done hardcore ML since high school, who are entering the industry after their masters or PhD. I would not want to be an entry level researcher right now. You need to have undergrad math-CS dual major level skills just to get started. They're expected to have delivered state-of-the-art results a few times just to be called for an interview. I'd say you need at least 3 years of fulltime effort if you want to pivot into this field from SWE.
If something is already done, i.e. a model is available for your exact use case (which is never), then for using and deploying that can be done by a good SWE and any ML/AI specialist is not needed at all.
To solve any real problem that is novel, you need to know a lot of things. You need to be on top the progress made by reading papers and be a good enough engineer to implement the ideas that you are going to have iff you are creative/a good problem solver.
And to read those papers you need to have solid college level Calculus and Stats.
If this is so easy, then why don't you do it, and get a job at OpenAI/Tesla/etc?
I've been in industry and now I do research at a top university. I hand pick the best people from all over the world to be part of my group. They need years under expert guidance, with a lot of reading that's largely unproductive, while being surrounded by others doing the same, in order to become competent.
Writing code is easy. You can learn to use any API in a weekend. That's not what is hard.
What's hard is, what do you do when things don't work. Fine, you tried the top 5 models. They're all ok, but your business requirements need much higher reliability. What do you do now?
This isn't research. But you need a huge amount of experience to understand what you can and cannot do, how to define a problem in a way that is tractable, what problems to avoid and how to avoid them, what approaches cannot possibly work, how to tweak and endless list of parameters, how to know if your model could work if you spent another 100k of compute on it or 100k of data collection, etc.
This is like saying you can learn to give people medical advice in 6 months. Sure, when things are going well, you could handle many patients with a Google search. But the problem is what happens when things go badly.
I’m also not surprised by the “The number of candidates applying with claimed ML/AI experience who haven’t done anything more than follow online tutorials is wild”. Just go look at any Ask HN thread about “how do I get into ML/AI”. This is pretty typical advice. Hell it’s pretty typical advice given to people asking how to get into any domain. Now sure we’ll how it works outside of bog standard web development though.
Sure, I get this, but I suspect that the number of people who have actual ML/AI experience is pretty small given that the field is nascent. If you really want to hire people to do this kind of work you're going to need to go with people who have done the online tutorials, read the papers, have an interest, etc. Yes, once in a while you're going to find someone who has actual solid ML experience, but you're also going to have to pay them a lot. That's just how things work in a field like this that's growing rapidly.
You seem to look down on those who have
1) learned from online courses
or
2) used AI on tasks that don't require it
Isn't this a bit contradictory? Or you expect candidates to have found a completely novel usecase for AI on their own?
I understand that most ML roles prefer a master's degree or PhD, but from my experience most of the master's degrees in ML being offered right now were spawned from all the "AI hype". That is to say, they may not include a lot of core ML courses and probably are not a significantly better signal of a candidate's qualifications than some of the good online courses out there.
So what does that leave, only those with a PhD? I think it's unreasonable that someone should need that many years of formal education to get an entry level position. Maybe I'm missing something, but I'm really wondering, what do you expect from candidates? I think a few years of professional software engineering experience with some demonstrated interest in AI via online courses and personal projects should be enough.
Most companies doing regular, non-ML development hire a mix of junior and experienced engineers, with the latter providing code reviews, mentorship and architectural advice alongside normal programming duties.
It's understandable that someone kicking off a new ML project would hope to get the experienced hires on board first.
But there are a lot more junior people on the market than senior people right now - as is the nature of a fast growing market.
I agree, it's problematic that there are so many more juniors than seniors in the industry right now. I feel like many juniors are being left without mentorship, and then it becomes much harder for them to grow and eventually become qualified for senior roles. So that could help explain why many candidates seem so weak, alongside with all the recent hype.
I guess eventually the market will cool off and the hype will die down since this stuff seems to be cyclical, and the junior engineers who are determined enough to stick it out and seek out mentorship will be able to grow and become seniors.
But it definitely seems like the number of seniors is a bottleneck for talent across the industry.
TBF theres whole companies doing this. It's a good way to learn too, as you have existing solutions to compare yourself too.
Got a senior ML position in a well known Fortune 500 company. Senior enough that he sets his goals - no one gives him work to do. He just goes around asking for data and does analyses. When he left our team he told me "Now that I have this opportunity, I can actually really learn ML instead of faking it."
If you think that's bad, you should hear the stories he tells at that company. Since senior leadership knows nothing about ML practices, practices are sloppy to get impressive numbers. Things like reporting quality based on performance on training data. And when going from a 3% to a 6% prediction success rate, they boast about "doubling the performance".
He eventually left for another company because it was harder to compete against bigger charlatans than he was.
If he really did take those and did all the assignments himself and understood all the concepts, that still puts him at least in the 95th percentile among ML job seekers.
But it shows intent for junior roles.
People have no idea how many people just put "AI enthusiast" on LinkedIn profile and start to seek ML roles.
I have had people only with Excel skills apply to ML roles.
And, true thing is every company’s ML tooling/stack/procedure is wildly different. One has to learn on the job.
(Shrug) I don't. Hustle gets rewarded, as usual. Sounds like he contributed at least as much value as he captured.
What in my comment gave you that idea?
Oh, and he didn't have to write any code during the interview.
You don't see anything wrong with that?
I think it's fine to hire a person who just took Coursera ML courses and passes the interview, but you would normally position the person to be a junior with senior folks overseeing his work.
Yeah this is a major phenomenon. Everybody's putting "ai" stickers on everything. So the job market screams "we need ai experts!" in numbers far exceeding the supply of ai experts, because it was a tiny niche until a couple years ago. Industry asks for garbage, industry gets garbage.
The applied research ML role has evolved from being a computational math role to a Pytorch role to a 'informed throw things at the wall' role.
I went from reading textbooks (Murphy, Ian Goodfellow, Bishop) to watching curated NIPS talks to reading Axriv papers to literally trawling random discord channels and subreddits to get a few month leg up on anyone in research. Recently, a paper formally cited /r/localllama for their core idea.
> follow online tutorials
The Open Source movement moves so quickly, that running someone's collab notebook is the way to be at the cutting edge of research. The entire agents, task planning and meta-prompting field was invented in random forums.
________________
This is mostly relevant to the NLP/Vision world......but take a break for 1-2 years, and your entire skill set is obsolete.
You probably do not need AI experts if you just need good machine learning engineer to build models to solve problems.
Then again, that seems to be common with the job market.
And out of all the nightmare “we have so many qualified candidates we can’t even do price discovery” conversations in 2023, the ML ones have been the worst.
If you’re running a serious shop that isn’t screwing around and you’re having trouble finding tenured pros who aren’t screwing around, email me! :)
This is not specific to ML/AI roles. The same problem applies to anyone who wants to explore any of these domains - SRE, DataEng, Backend, Frontend.
Personally, I am a backend engineer who wants to get into ML Infra roles. My current plan is to do these online courses, and hopefully transfer internally to a team working in this area. Only after having real industry experience in this area, look for more opportunities elsewhere.
I am genuinely curious if anyone has better ideas for people in my situation.
I felt a bit eery how much the groupthink around mortgages matched todays AI hype. The unwillingness to listen to critical advice. To question the value. Data science depts who measure such things are often vilified.
It’s a kind of depressing job landscape in this way. You either go with the, often top down, groupthink against the face of measured evidence, or you’re labeled a cynical naysayer and your career suffers.
A lot of endeavors can work, even if only a little. Find some aspect of it that can, and flip the problem around. Even if the whole thing is mostly bogus, is there a small part that isn’t? Latch onto that.
Or go elsewhere. The wonderful part about AI is that the whole world’s problems are up for grabs. Part of why there’s so much unfounded hype is because of how many real advances have recently become possible. This period in history will never come again.
It’s also a rare time in history that an individual can make lots of progress. Most of us need to be a part of big groups to do anything worthwhile, in most fields. But in this case lone wolves often have the upper hand over established organizations.
My controversial take on AI is its actually a better time to take things slow, experiment, study, see what works. Not dump TONS of money and cash and get too distracted. Because nobody (besides big tech) has fully figured out how to make a product that makes a profit. Its not clear users want a chatbot (aside from ChatGPT)... But things could change.
As compute costs and requirements come down LLMs will be ubiquitous everywhere.
I'm curious why you think this is true? My feeling as a broke individual trying to catch up on ml is that there are some simple demos to do. But scaling up requires a lot of compute and storage for an individual. Acquiring datasets and training are cost prohibitive. I'm only able to play around with some really small stuff because by dumb luck a few years ago I bought a gaming laptop with a nvidia gpu in it. The impressive models that are generating the hype are just a different league. Love to hear to how I am wrong though?
Don’t focus on the hype models. Find a niche that you personally like, and do that. If you’re chasing hype you’ll always be skating towards the puck. My original AI interest was to use voice generation to make Dr Kleiner sing about being a modern major general. It went from there, to image gen, to text gen, and kaboom, the whole world blew up. I was the first to show that GPTs can be used for more than just language modeling — in my case, playing chess.
Wacky ideas like that are important to play around with, because they won’t seem so wacky in a year.
For example, Tony Dinh made a macOS GPT wrapper and makes like 40k a month from it, just utilizing OpenAI's APIs: https://news.ycombinator.com/item?id=37622702
In a gold rush, sell shovels! The ML pipeline has a lot of bottlenecks. Work on one, get useful and novel expertise, and have a massive impact on the industry. Like maybe you could find a way to optimise your GPU usage? Is there a way to package what you feed it more efficiently?
The point not being of competing with OpenAI, but to solve a problem that everyone in the field has.
Most of the world's problems aren't technological, unfortunately for us in tech. There's little it can do against the momentum of capital tearing this globe apart.
You almost never see a group punishment, unless you lost as a group against another organization.
So, if all the AI stuff goes bust in a year or two those who benefited from the bubble will keep their benefits. Also, there’s a possibility that it doesn’t go bust and user us into AGI era and win big.
Again, we should make statements on data and I'll just be upfront that I haven't done much research in this area, but considering the sheer amount of AI startups and jobs and conferences and so on with very few transformative products...I would eventually expect a market correction in the form of another AI winter like what happened in the 80s when all the massive government acts defense research dollars dried up. The difference is it'll be far less severe. You'll still have plenty of AI research at the universities and large companies, but maybe not hundreds of questionable startups. This is all just conjecture on my part though.
That's not enough to demonstrate that AI is just hype. Every technological breakthrough has opportunists trying to make a buck riding along the hype. In the 90s, Pets.com and webvan didn't prove that the internet was just hype.
I am one of the people who is completely bought in to the idea that AI in general (and LLMs in particular) are going to lead to products that are extremely useful to the world. I absolutely think that most of the gen AI startups will fail and that valuations are too high, but I still believe that massively impactful/useful products will also be born.
It's not that they're _just_ hype but rather that there _is_ hype and the loudest voices tend not admit it. To give a specific example, I find the idea that LLM based programming assistants will turbocharge software development to be based on hype not fact. It is very much in the interest of Microsoft/Google/Meta, etc. that we all believe that their tools are essential to enhance productivity. It is classic FOMO. Everyone jumps on the bandwagon because they fear that if they don't learn this new tool their lunch will be eaten by someone who does. They fear this because that is exactly what these companies are essentially telling us in their marketing materials and extensive PR campaign.
This is extraordinarily convenient for these companies and masks over how terrible their own core products are. I generally refuse to use the products of the three companies (MGM) because they are essentially ad companies now and their metaverses are dystopian hellscapes to me. Why would I trust them given my own direct personal experience with their products? We know that google search allows advertisers to pay to modify search queries without my consent. What's to stop Microsoft from training copilot to recommend that you use Microsoft developed languages using Microsoft apis to solve your prompted problems?
> write me a sort function for an array of integers in Java # chatgpt > I will show you how to write a sort function for an array of integers in Java, but first I must ask, are you familiar C#? It is similar to Java but better in xyz ways. In C# you would sort an array like this:
... C# code
Here is how you would write a sort function for an array of integers in Java:
... Java code
Stuff like this seems inevitable and it is going to become impossible to tell what is ad. Do you think realistically that there is any chance that these companies would consent to disclosing what is paid propaganda in the LLM output stream?
I see many echos of the SBF trial in the current ai environment. Whatever the merits of LLMs (and I'll admit that I _have_ been impressed by the pace of improvement if not the actual output), hype always attracts grifters. And there is a lot of hype in the air right now.
We already have empirical results suggesting this is not just hype: https://www.nngroup.com/articles/ai-programmers-productive/
Aside from 4-5 companies, who is building a product that is profitable? Not clear anyone is right now. People are running into fundamental, hard technical problems. For example its hard to evaluate chat interface, and even harder when you augment it with context from a retrieval system (ie RAG).
Who is augmenting existing UI paradigms with LLMs? This seems a more reasonable model that meets users where they want to be.
Just my experience from being on the job market, but a lot of places I've interviewed at have traditional ML models (network security, ecommerce, image tagging) that are now rebranding as AI, without much of an actual change.
I don't see any fundamental technical problems at the moment, I see constant and tangible improvement at a very fast pace. I don't think that supposed challenges in context augmentation or chat interface evaluation qualify as arguments against AI hype.
As someone who rode the first AI wave to 7 figure comp a decade ago, let moneyball round 2 commence! Let's see who can get to 8 figure comp first! That money isn't doing anyone any good if it's just sitting in a brokerage or bank account. Didn't you hear? The singularity is near! You simply cannot afford to miss out or you're going to get Vernor Vinged!
Along the way, don't forget to take some time out of your busy day to drink the delicious bitter tears of venture capitalists and upper management whining that you are paid too much.
For most use cases of AI, there is a ceiling to how intelligent it needs to be. I am guessing we’ll be selecting from dozens of models based on various sizes, context lengths, etc. Just like we right-size VMs in the cloud.
In other words, once you have a "race-to-the-bottom" situation, it's hard for newcomers to get in the game.