Understanding Deep Learning
udlbook.github.io
udlbook.github.io
Both perspectives are correct. The field is bifurcating into two different skill sets: ML engineer and ML scientist (or researcher).
It's great to have both types on a team. The scientists will be too slow; the engineers will bound ahead trying out various APIs and open-source models. But when they hit a roadblock or need to adapt an algorithm many engineers will stumble. They need an R&D mindset that is quite alien to many of them.
This is when an AI scientists become essential.
The people who create the models and the people that use them.
The people who create the programming languages and the people that use them.
Whereas it's unlikely in most programming jobs you would need to do any research into programming language design.
The point the commenter is making is that both schools of thought in the comments are valuable and unless you perform both roles, i.e. an engineer who is familiar with the scientific foundations, both are symbiotic and not in contention.
It's almost self-exploratory that when you hit a roadblock in practice you go back to foundations, and good people should aim to do both. In that case I don't see where ML engineer/scientist bifurcation comes from except for some to feel good about themselves
There is a need for people who are able to build using available tools, but who don't have an interest in the theory or foundations of the field. It's a valuable mindset and nothing in my original comment suggested otherwise.
It's also pretty clear that many comments on this post divide into the two mindsets I've described.
My experience is the other way around.
People underestimate how powerful building systems is and how most of the problems worth solving are boring and require out-of-the-box techniques.
During the last decade, I was in some teams and I noticed the same pattern: The company has some extra budget and "believes" that their problem is exceptional.
Then goes and hires some PhDs Data scientists with some publications but only know R and are fresh from some Python bootcamps.
After 3 months, or this new team no much was done, tons of Jupyter notebooks around but no code in production, and some of them did not even have an environment to do experimentation.
The business problem is still not solved. The company realizes that having a lot of Data Scientists not not so many Data/ML Enginers means that they are (a) blocked to do pushing something to production or (b) are creating a death star of data pipelines + algorithms + infra (spending 70% more of resources due to lack of knowledge on Python).
The project gets delayed. Some people become impatient.
Now you have a solid USD 2.5 million/year team that is not capable of delivering a proof of concept due to the fact that people cannot do the serving via Batch or via REST API.
The company lost momentum, competitors moved fast. They released an imperfect solution, but a solution ahead, and they have users on it and they are enhancing.
Frustration kicks in, and PMs and Eng Managers fight about accountability. VP of Product and Engineering wants heads in a silver plate.
Some PhDs get fired and go to be teachers in some local university.
Fin.
https://nonint.com/2023/06/10/the-it-in-ai-models-is-the-dat...
Unlike other areas of ML, the nature of deep learning is such that its parts are interoperable. You could use a transformer with a CNN if you wish. Also, deep learning enables you to do machine learning on any type of data, text, images, video, audio. Finally, it can naturally scale computationally.
As someone pretty involved in the field, I lament that LLMs are turning people away from ML and deep learning, and following the misconceptions that there’s no reason to do it anymore. Large algorithms are expensive to run, have slow throughput and are still generally poorer performing than purpose built models. They’re not even that easy to use for a lot of tasks, in comparison to encoder networks.
I’m biased, but I think it’s one of the most fun things to learn in computing. And if you have a good idea, you can still build state of the art things with a regular gpu at your house. You just have to find a niche that isn’t getting the attention that LLMs are ;)
The whole thing is essentially curve fitting. The ML field is essentially an art more than a science and it's all about tricks and intuitions on different ways of getting that best fit curve.
From this angle the whole field got way less interesting. The field has nothing deeper or more insightful to offer beyond this concept of curve fitting.
I think there are some really cool problems, such as:
1. Is synthetic data viable for training?
2. How do you make deep learning agents that can do task planning and introspection in complex environments?
3. How do we efficiently build memory and data lookup into AI agents? And is this better/worse than making longer context windows?From a technical standpoint, it’s not correct analogy either, because it assumes you have a curve to fit. What curve is language? What’s curve is images? No answer, because there isn’t one. Deep learning is about modeling complex behaviors, not curve fitting. Images and language for instance are based in social and cultural patterns and not intrinsic curves to be fit.
At best, it’s an imprecise statement. But I’d disagree entirely.
IOW: to me, fitting a generalized linear model is very different than fitting a convolutional network.
In terms of continued relevance - "deep learning", meaning, dense neural nets trained to optimize a particular function, haven't fundamentally changed in practice in ~15 years (and much longer than that in theory), and are still way more important and broadly used than the OpenAI stuff for most purposes. Anything that involves numerical estimation (e.g., ad optimization, financial modeling) is not going to use LLMs, it's going to use a purpose-built model as part of a larger system. The interface of "put numbers in, get number[s] out" is more explainable, easier to integrate with the rest of your software stack, and more measurable. It has error bars that are understandable and occasionally even consistent. It has a controllable interface that won't suddenly decide to blurt corporate secrets or forget how to serialize JSON. And it has much, much lower latency and cost - any time you're trying to render a web page in under 100ms or run an optimization over millions of options, generative AI just isn't a practical option (and is unlikely to become one, IMO).
I don't have a significant math or theoretical ML background, but I've spent most of the last 10 years working side by side with ML experts on infra, data pipelines, and monitoring. I'm not sure I could integrate the sigmoid off the top of my head, but that's not what's important - I've done it once, enough to have some idea how the function behaves, and I know how to reason about it as a black box component.
Which video is this?
His series on making a GPT from scratch is also great for building intuition specifically about text-based generative AI, with an audience of software developers.
So yes if you are prompt engineering, and wondering why X works and sometimes it doesn't, and why any of this works at all, it is good to study a bit.
After a glance, looks like too much for one book. Probably it was compressed with the assumption that reader already knows quite a lot. In other words it's not an easy reading.
OpenAI is for example not interested in developing small embedded neural networks that run on a sensor chip that real-time detects specific molecules in air.
For the impatient, look into slide #123. Essentially, the recommendations are Murphy, Gelman, Barber, and Deisenroth.
Note these slides have a Bayesian bias. In spite of that, Murphy is a great DL book. Besides, going through GLMs is a great way to get into DL.
Joking aside, these slides are excellent! Is there an associated video or course that they were a part of?
Fun facts, the infamous Attention paper is closing in to reach the 10K citations, and it should reach this milestone by the end of this year. It's probably the fastest paper ever to reach this significant milestone. Any deep learning book written before the Attention paper should be considered out of date, and needs updating. The situation is not unlike an outdated Physics textbook with Newton's laws but devoid of the infamous Einstein's equation of energy equivalence.
I'm worried that I'm starting a journey that requires a Master's or PhD.
... which will be easier it you have a solid grasp of the foundations of the field. If you only ever focus on the "latest shiny" you'll be lost and left floundering when the landscape changes out from underneath you.
Machine learning platforms become obsolete.
Machine learning algorithms and ideas don't. If learning SVN or Naive Bayes did not teach you things that are useful today, you didn't learn anything.
It's almost like arguing that everything you learned as a Java developer is completely useless when a new programming language replaces it.
So you really did not learn them.
There is nothing wrong with being user. You don't have to know how compilers work to use compiler. But then you should not say you understand compilers.
In the same way, you probably would benefit from a book "Using deep learning", not "Understanding deep learning".
Exactly my point. You are so into user perspective that you think you are arguing against me.
Nobody deploys a textbook algorithm because everyone knows textbooks algorithms and there are no advantages. So, no, there is real value in learning the fundamentals, dear founder.
To attempt answering this question, we can look at LLMs as an analogy. If you include code in the training set for an LLM, it also makes the LLM better at non-coding tasks, suggesting that sometimes learning something makes you also better at other things. I'm not saying the same necessarily applies for learning these "old school" AI techniques, but it's a decently analogy at least.
Also, being an "expert in LSTM" is like being an "expert in HTTP/1.1" or "knowing a lot about Java 8". It's not knowledge or a skill that stands on its own. An expert in HTTP/1.1 is probably also very knowledge about web serving or networking or backend development. HTTP/2 being invented doesn't obsolete the knowledge at all. And that knowledge of HTTP/1.1 would certainly come in handy if you were trying to research or design something like a new protocol, just as knowledge of LSTMs could provide a lot of value for those looking for the next breakthrough in stateful models.
If it became obsolete, then y'all were doing the new shiny.
The fundamentals don't really change. There are several different streams in the field, and there are many, many algorithms with good staying power in use. Of course, you can upgrade some if you like, but chase the white rabbit forever, and all you'll get is a handful of fluff.
Who is the author ?
Have they published anything else highly rated ?
Are there good reviews from people that know what they're talking about?
Are there good reviews from students that don't know anything ?
The entire pdf is available as a free download on that page. First link at the top.
https://github.com/udlbook/udlbook/releases/download/v1.16/U...
based on a table of contents? You can download the draft of Chapters 1-21 (500+ pages) from the linked site.
Who is the author ? Simon J. D. Prince is Honorary Professor of Computer Science at the University of Bath and author of Computer Vision: Models, Learning and Inference. A research scientist specializing in artificial intelligence and deep learning, he has led teams of research scientists in academia and industry at Anthropics Technologies Ltd, Borealis AI, and elsewhere.
Have they published anything else highly rated ? Author of >50 peer reviewed publications in top tier conferences (CVPR, ICCV, SIGGRAPH etc.) https://scholar.google.com/citations?user=fjm67xYAAAAJ&hl=en
Are there good reviews [...] The book has not been published, this is literally a free draft that you are looking at. The book is listed on Amazon as a pre-order for 85USD.
They are less popular, and less explored. But an interesting route ahead.
Which suggests two obvious paths forward:
1. Don't bother learning / using RNN's
2. Co-develop new hardware / new RNN architectures that work together to provide great performance per unit of price.
Now of course nobody is saying (well, I am not saying) that (2) would be easy... or even necessarily possible. But somebody should at least be intrigued by the idea. And in the world we live in today where FPGA's and other devices make it easier than ever to experiment with custom hardware architectures... it might be worth taking a stab at it.
I have a discord https://discord.cofunctional.ai