Neural Networks: Zero to Hero
karpathy.ai
karpathy.ai
Also, the way he builds up everything is magnificent. Starting from basic python classes, to derivatives and gradient descent, to micrograd [2] and then from a bigram counting model [3] to makemore [4] and nanoGPT [5]
[1]: https://www.foundersandcoders.com/ml
[2]: https://github.com/karpathy/micrograd
[3]: https://github.com/karpathy/randomfun/blob/master/lectures/m...
I’ve been simply watching them on a palm from a hammock and I’m worried I’m not getting the full experience.
Watching only is much nicer for entertainment.
There is simulated environment in which all teams of the cohort receive millions of requests per day (and hundreds of thousands of users and items) and you have to build out your infrastructure on an EC2 instance, build a basic model, and then iteratively improve on it. Imagine a simulated facebook/youtube/tiktok-style system where you aim for the best uptime and the best recommendations!
My email address is first part of my username (before the “-“) at blueheart dot io.
The only aspect I could see being non-ideal for some is that it uses some Python-specific cleverness/advanced syntax and semantics (__call__(), list comprehensions with two for's, **kwargs, __add__, __repr__, subclasses, (nested) functions as variables etc.), but if you are familiar with these it might seem more compact and elegant as well.
But this does not remove any credit to Andrej’s class.
It was very satisfying to learn how transformers worked, to finally be able to turn the obscure glyphs of the research papers into real code, but I think transformers are too big for what I can do on my own computer. The author mentioned that the toy transformer he was building in the final video took 15 minutes to train on his A100 GPU (a $10,000 GPU), and the results weren't even that good; the transformer was spelling words correctly using character level tokens, I guess that's something, but it's not GTP4.
Even so, there were a lot of good tips to pick up along the way. This is a great series that I'm thankful to have. The "Backprop Ninja" video was hard work, you manually calculate the gradients and then compare your calculations against PyTorch. It's great to have instant feedback telling you whether your gradients are correct or not.
These are available to rent per hour at much lower costs. The author mentions this in the video description.
$1.29 per hour for a 40gb a100 apparently
https://lambdalabs.com/service/gpu-cloud#pricing
$1.10 per hour
With the original hyperparameters, it was 30-60 minutes, with a pruned down network and adjusted hyperparameters, about 6 minutes, and a variety of optimizations beyond that to bring it down.
If you want the nano-GPT basically feature-identical (but pruned down) version, 0.0.0 at ~6 minutes or so is your best bet.
You can get A100s cheaply and securely through Colab or LambdaLabs.
Thanks karpathy!
https://deepai.org/machine-learning-glossary-and-terms/logit....
The inverse logit or logistic function takes log odds and transforms them back into probabilities.
Most machine learning relies heavily on manipulating probabilities, but since probabilities are not linear, the logit/logistic transformations become essential to correctly modeling complex problems involving probabilities.
For example in many classification problems you get a 1D vector of logits from the final layer, you apply softmax to normalise, then argmax to extract the predicted class. It extends to other tasks like semantic segmentation (predict pixel classes) where the "logit" output is the same size as the image with a channel for each class and you apply the same process to get a single channel image with class-per-pixel.
Here's a nice explanation: https://stackoverflow.com/a/66804099/395457
Before going through the softmax layer, the logits will be small numbers around 0, probably. Something like: [2.89, -4.53, 0.24, -1.556, 0.57]. Logits like this are natural outputs of a neural network, because they can be any real number and everything will still work.
The logits become percentage as follows:
julia> x = [2.89, -4.53, 0.24, -1.556, 0.57]
5-element Vector{Float64}:
2.89
-4.53
0.24
-1.556
0.57
julia> x = e.^x
5-element Vector{Float64}:
17.993309601550315
0.010780676072743085
1.2712491503214047
0.2109782988178321
1.768267051433735
julia> x / sum(x)
5-element Vector{Float64}:
0.8465613320288766
0.0005072164987105474
0.05981058503789324
0.009926248902037537
0.08319461753248213
Logits is an overloaded term though, and means different things in different contexts.Generally I'm a huge fan of not getting too caught up in theory before diving into practice, but I'm seeing multiple responses to this comment without a single mention of "log odds".
The logit function transforms probabilities into the log of the odds (ln P(X)/(1-P(X)), which is important because it makes probabilities linear, which they are not in their standard [0,1] form. It's the foundation of logistic regression, which is, despite much misinformation, quite literally linear regression with a transformed target.
The logistic function is the inverse of the logit: it turns log odds values into probabilities once again. Logistic regression actually transforms the model not the target (most of the time) because the labels 0 and 1 are negative and positive infininity which can't be handled by linear regression (so we transform the model using the inverse instead).
I don't think I can stress enough how important is its to really understand logistic regression (which is also the basic perceptron) before diving into neural networks (which are really just an extension of logistic regression).
Hopefully, I'll manage to get further with this course.
It's an old site and guide, but probably still the easiest to understand if you're coming from a programming background.
I can even recommend his interview with Lex Fridman.
Let's build GPT: from scratch, in code, spelled out: https://www.youtube.com/watch?v=kCc8FmEb1nY
edit: Python really was/is made for this numbers/calculation/visualization thing. Kinda kicking myself now for not investing more in it and sticking with PHP, although PHP has its merits when building different things, Python is a beast with numbers.
Also using ChatGTP to ask questions where I don't get something.
Wow what a time we live in to learn things.
https://uvadlc-notebooks.readthedocs.io/en/latest/index.html
Tutorial 6 covers transformers.
There are several open courses, online.
If you want, you can also start with the first lessons of course.fast.ai
Having these disconnected pieces of information with no clear link to one another feels like a lot of noise to me.
His explanation of attention is the most accessible I have ever seen.