Intro to Large Language Models [Video]
youtube.com
youtube.com
I was thinking of this other video he had published early this year: "Let's build GPT: from scratch, in code, spelled out" https://www.youtube.com/watch?v=kCc8FmEb1nY
Karpathy is generally well-reputed as a good tutor, especially for complex topics in AI / ML.
System 1 thinking: Fast automatic thinking and rapid decisions. For example is when someone ask you 2 + 2, you don't think. You just reply quickly instantly. LLMs currently only have system 1 thinking.
System 2 thinking: Rational slow thinking to make complex decisions. For example when someone ask you 17 x 24 you think slowly and rationally to multiply. This kind of thinking is a major component we need for AGI. Current rumor from OpenAI about so called "Q*" algorithm could be something related to system 2 thinking (Just speculation at this point)
System 1 thinking: patterns.
System 2 thinking: logic.
For example:
System 1 thinking: Does x sentence sound like a correct English sentence.
System 2 thinking: Verify x sentence is a correct English sentence by using grammar rules.
Someone fluent in English can form correct English sentences using only system 1 thinking, while someone that has just started learning English must think about grammar rules (using system 2 thinking) to do it.
For example I've noticed that a lot of the time when I ask ChatGPT a coding question it might get 90% of the answer. When I tell it what to fix and/or add, it usually gets the answer. I wonder if they're using these refined answers to fine-tune those original prompts.
I wonder how the LLM interacts with other software like the calculator or Python interpreter. It would be great if this were modular so that the LLM OS could be more like Unix than Windows which is what OpenAI seems to be trying to emulate.
Ultimately though it seems to me like AGI is fairly straightforward from here. Just train on more quality data - in particular enabling the machine to generate this training data, increase parameter size, and the LLM just gets better and better. Seems like we don't even need any new major breakthroughs to create something resembling AGI.
There's something enlightening in hands-on learning without using metaphors. He even opens the code of production grade tools to show you how exactly the concepts he explained and build together are actually implemented IRL.
This is a style of teaching that clicks with me. I don't learn well with metaphors and high abstractions and find it magical to remove the magic of amazing things and bring it down to easy to reason pieces which can create a complex structure with composition so you can just disregard the complexity as a separate thing of the core.
An aside... incredibly, it looks like he recorded in one cut from his hotel room.
If you question the response and check it against responses about related questions, how much does it align with those?
This is what we humans do, too. It’s also what we do in science. It’s things not “adding up” that tells us where we must improve.
Highly recommend!
https://youtu.be/XfpMkf4rD6E?si=1_EmuYDFfi7RNEhz
This video is the best for learning attention, specifically where he explains:
Think of attention like a directed graph of vectors passing messages to each other
Keys are what other tokens are communicating to you,
Queries are what you are interested in,
and Values are what you are projecting out yourself.
When you matrix multiply the queries x keys(transposed), you measure the interestingness or affinity between the two.
Meanwhile a guy somewhere in Africa adjusting an answer probably stating that humans can do photosynthesis: bruh