TL;DR of Deep Dive into LLMs Like ChatGPT by Andrej Karpathy
anfalmushtaq.com
anfalmushtaq.com
I am going through the video myself -- roughly halfway through -- and have a fw things to bring up.
Here they are now that we have a fresh opportunity to discuss:
1 - MATH and LLMs
I am curious why many of the examples Andrej chose to pose to the LLM were "computational" questions -- for instance "what is 2+2" or some numerical puzzles that needed algebraic thinking and then some addition/subtraction/multiplication (example at 1:50 mins about buying Apples and Oranges).
I can understand these abilities of LLMs are becoming powerful and useful too -- but in my mind these are not the "basic" abilities of a next token predictor.
I would have appreciated a more clear distinction of prompts that showcase core LLM ability -- to generate text that is acceptable as generally grammatically correct, based in facts and context, without necessarily needing the ability of a working memory / assigning values to algebraic variables / doing arithmetic etc.
If there are any good references to discussion on the mathematical abilities of LLMs and the wisdom of trying to make them do math -- versus simply recognizing when a math is needed and generating the necessary python/expressions and let the tools handle it.
2 - META
While Andrej briefly acknowledges the "meta" situation where LLMs are being used to create training data for the training of and judge the outputs of newer LLMs ... there is not much discussion on that here.
There are just many more examples of how LLMs are used to prepare mitigations for hallucinations by preparing Q&A training sets with "correct" answers etc
I am curious to know more about the limitations / perils of using LLMs to train/evaluate other LLMs.
I kind of feel that this is a bit like the Manhattan project and atomic weapons -- in that early results and advances are being looped back immediately into the development of more powerful technology. (A smaller fission charge at the core of a larger fusion weapon -- to be very loose with analogies)
<I am sure I will have a few more questions as I go through the rets of the video and digest it>
Somewhere in the video he says that LLMs have expert (only slightly fuzzy) knowledge about a lot of topics, but fail with simple math questions. Many non-technical people anthropomorphize LLMs and don't know that they can't think or calculate like a real calculator. LLMs compute tokens and you can improve the performance, if you don't put too much computation into a single result token.
I think it's an excellent example to show the capabilities and limits of LLMs. For softer topics, you can argue a lot more about what's considered to be right or wrong. With Math, you have a single correct answer that can be evaluated and people assume that computers are good at computer things, such as calculating numbers, even though LLMs actually aren't good at this.
The takeaway is: Prompting and "computational complexity per token" matter and if you understand how it works for math, you probably understand how it works for softer things like answers about law or whatever.
there is only one algebraic approach to solving something like 2+2 and that is counting! 2+2 = (((0 + 1) + 1) + 1) + 1). but llms are infamously bad at counting. which is why 2+2 isn't an algebraic problem to an llm. it's pattern matching or linguistic reasoning token by token.
https://x.com/yuntiandeng/status/1889704768135905332
Is this a consequence of the fact that "multiplication tables" For kindergarteners are available online (in training data) abundantly ... typically up to 12 times or 13 times table as plain text ?
Also for the second point, check later in the video when he talks about RL and (simulated) RLHF - he gets into the feedback loops of models training each other and the collapse that follows.
At the extreme, there is this paper on ‘inbred LLMs’: https://www.nature.com/articles/s41586-024-07566-y
the entropy goes up in such case (up means less information). The result will be as if someone recompressed mpeg with another lossy compression. You can sometimes see the results on the internet.
- Extract a snippet of training data.
- Generate a factual question about it using Llama 3.
- Have Llama 3 generate an answer.
- Score the response against the original data.
- If incorrect, train the model to recognize and refuse incorrect responses.
In a way this is obvious in hindsight, but it goes against ML engineers natural tendency when detecting a wrong answer: Teaching the model the right answer.Instead of teaching the model to recognize what it doesn't know, why not teach it using those same examples? Of course the idea is to "connect the unused uncertainty neuron", which makes sense for out-of-context generalization. But we can at least appreciate why this wasn't an obvious thing to do for generation 1 LLMs.
But the answer space for LLMs is infinite and unbounded. So, no effort will be complete and you will always end up with the question of how to deal with uncertainty.
But I admit this is a bit of hindsight 20/20.
Edit: concretely, we can presume that OpenAI didn't specifically train ChatGPT to know that "Orson Kovacs" isn't a famous person, right? That's all I'm saying here - that they trained it how to say it doesn't know things, and it took care of the rest.
Hard to buy.
If a machine makes a mistake, it's because it was configured wrong or because of wear and tear, solar flares or some quake or some manufacturing defect in a part. If a learning machine makes a mistake, it's because it's learning has not extended it's rule set to cover that matrix/mistake/pattern, yet; and so it includes that mistake/matrix and other mistakes, analyses for patterns and then creates mistakes that fall into that pattern. Later doing that in a rolling release or canine kind of way and even later learning machines will do it all live, synchronous to their concurrent actions.
But yeah, thinking about that, I see why ML engineers wouldn't get there from scratch. It's a rhythm, after all, an epiphany about or realization of how ones dog, ones brain works, learned and then coded step by step. And there is, of course the variety of how people learn and "realize".
Someone has to show us the work of those savant programmers/engineers I still haven't seen a documentary of.
Also how can the training of LLMs be parallelized when updating parameters are sequential? Sure we can train on several samples simultaneously, but the parameter updates are with respect to the first step.
(Hence the analogy to training AlphaGo, wherein you take a model that sometimes wins games, and then play a bunch of games while reinforcing the cases where it won, so that it evolves its own ways of winning more often.)
In the LLM case you have to have an already capable model to do RL. Also I feel like the problem selection part is important to make sure it's not too hard. So there's still much labor involved.
https://medium.com/@sahin.samia/the-math-behind-deepseek-a-d...
Destroying the copies they took will be what the courts ordered, but the data will still be there.
See The Open Source AI Definition from OSI: https://opensource.org/ai
With LLMs, the list doesn't even have to be kept up to date, nor the links alive (though publishing content hashes would go a long way here). It's not like you can get an identical copy of a model built anyway, there's too much randomness at every stage in the process. But, as long as the details of cleanup and training are also open, a list of training material used would suffice - people would fetch parts of it, substitute other parts with equivalents that are open/unlicensed/available, add new sources of their own, and the resulting model should have similar characteristics to the OG one who we could, now, call "open source".
It's just not literally labelled so because of obvious reasons.
https://www.reddit.com/r/LocalLLaMA/comments/1ilsfb1/comment...
I recommend reading the actual Open Source AI Definition[1] and the FAQ[2]. There's also the whitepaper[3] that goes into much more detail about the state of affairs.
[1]: https://opensource.org/ai/open-source-ai-definition
[2]: https://hackmd.io/@opensourceinitiative/osaid-faq#What-is-th...
[3]: https://opensource.org/wp-content/uploads/2025/02/2025-OSI-D...
the model reveals the architecture which is all you need to use/run/train it.
For an LLM you can finetune and enhance, distill and embed given just the model weights, the runtime, and a permissive license. Having more is better. Well written detailed model release papers help a lot. Training code and training data are a great bonus.
However, I find the purity contest a bit too dismissive of the great contributions to the AI dev ecosystem that Meta and Deepseek have brought us. Without these, there wouldn't be the open ecosystem we have today.
If I write a program, then obfuscate it and then release the obfuscated code under an open source license, would you consider it open source(I would)? That's kind of the case here, they are releasing the model weights under an open source license.
Personally, I think it's fine to shorten it to "open source model" instead of "a model with the weights released under an open source license". What I would object to is releasing model weights under a restrictive license and calling that open source.
I wouldn't. Most definitions of open source say something like "in the form used for editing". You can release a built binary under an unrestrictive license, but that does not mean that you've opened the source. It's literally the plain meaning of the words: the source, as in where the thing comes from, needs to be open for it be meaningful.
But that's also true for binaries, games are a good example of where people pushed this quite far. Based on what little experience I have in ML, I'd say it's about the same thing. Whereas an API is more akin to a piece of software you can't tinker with in any way.
Guess the bar is just lower in the LLM space :P
Much in the same way, no sane company will touch the legal nightmare of releasing LLM training data scraped from public websites. Even releasing the LLM alone might be infringement, there are literally court cases being fought over this right now.
And that's fine! It's still valuable to have access to the source code, even if the "batteries" aren't included. Of course, if you really want to call it an open source model you should include the source for the data scraping/cleaning stages too; then the only thing missing would be the compute time and risk of acquiring dubiously-legal inputs.
I personally prefer a taxonomy like:
* Open weights: you can download the artifact and run it locally, not just use it through an application like chatgpt or an API.
* Open source: the code that created the artifact is provided in the same format that the authors used to work on it.
* Open data: the dataset that the source code was used on is available for download.
All three of those could be individually licensed or released, for 8 possible combinations. In the analogy to games, they would correspond to the licenses on the retail binary, the source code of the game, and the original uncompressed art assets or Blender projects, respectively.
I agree that open source doesn't mean open assets, but neither does open assets mean open source. You could make a linguistic argument that the training data is part of the "source" of the model (as in, from whence it came), but in any case the point is moot because neither the training data nor the code is open.
The AI crowd doesn’t care much for licenses anyway.
Google dropped MHA self attention which was a major idea that they showed to work. OpenAI saw and built an empire on FeedForward attention models which are (compared to most alternatives) super stable at generation. DeepSeek showed evidence it's possible to further push these models and use effectively compression in the model design to pass around sufficient information for training. (Hence the Latent) They also did a lot of other cool stuff, but the main "core of the model" difference is this part...
Other than that, the biggest hurdle has been hardware. There's probably no way you could get kit from 2010 without even aes acceleration to evaluate most full-fat mhlffa models let alone train them. There's been a happy convergence of matrix acceleration on GPUs for gaming, graphical and high fidelity stimulation work. This combined with matrix based ML maths combined with high throughput memory advances means we can do what we're doing with llms now.
So inevitable outcome or happy convergence? That's for historians to decide imo. I think it's a bit of both.
Earlier models had huge bottlenecks in terms of information limits and precision. (Auto encoders Vs uNets for example) And LSTM are still semi unstable.
Why the attention design as posited by Google works so well is part the skip forward and part "now we have enough information and processing power to try this".
It's well motivated but from a first principles up do we expect this to work well, it's a bit less well understood still. And if you're good at that you'll likely get a job offer very quickly.
The smaller the input for the same quality the quicker/better/faster we can iterate so everyone is pushing to get the minimum viable training time of a decent llm down to allow both ChainOfThought to get cheaper as a concept and to allow for iteration and innovation.
As long as we live in the future aspoused by early OpenAI of huge models on huge GPUs we were going to stagnate. More GPU always means better in this game, but smaller faster models means you can do even more with even less. Now the major players see the innovation heading into the multi llm instance arena which is still dominated by who has the best training and hardware. But I expect to see disruption there too in time.
Once upon a time ....
Language modelling in general grew out of attempts to build grammars for natural languages, which then gave rise to statistical approaches to modelling languages based on "n-gram" models (use last n words to predict next word). This was all before modern neural networks.
Language modelling (pattern recognition) is a natural fit for neural networks, and in particular recurrent neural networks (RNNs) seemed like a good fit because they have a feedback loop allowing an arbitrarily long preceding context (not just last n words) to be used predicting the next word. However, in practice RNNs didn't work very well since they tended to forget older context in favor of more recent words. To address this "forgetting" problem, LSTMs were designed, which are a variety of RNN that explicitly retain state and learn what to retain and what to forget, and using LSTMs for language models was common before transformers.
While LSTMs were better able to control what part of their history to retain and forget, the next shortcoming to be addressed was that in natural language the next word doesn't depend uniformly on what came before, and can be better predicted by paying more attention to certain words that are more important in the sentence structure (subjects, verbs, etc) than others. This was addressed by adding an attention mechanism ("Bahdanau attention") that learnt to weight preceding words by varying amounts when predicting the next word.
While attention was an improvement, a major remaining problem with LSTMs was that they are inefficient to train due to their recurrent/sequential nature, which is a poor match for today's highly parallel hardware (GPUs, etc). This inefficiency was the motivation for the modern transformer architecture, described in the "Attention is all you need" paper.
The insight that gave rise to the transformer was that the structure of language is really as much parallel as it is sequential, which you can visualize with linguist's sentence parse trees where each branch of the tree is largely independent of other branches at the same level. This structure suggests that language can be understood by a hierarchy (levels of branches) of parallel processing whereby small localized regions of the sentence are analyzed and aggregated into ever larger regions. Both within and across regions (branches), the successful attention mechanism can be used ("Attention is all you need").
However, the idea of hierarchical parallel processing + attention didn't immediately give rise to the transformer architecture ... The researcher who's idea this was (Jakob Uszkoreit) had initially implemented it using some architecture that I've never seen described, and had not been able to get predictive/modelling performance to beat the LSTM+attention approach that it was hoping to replace. At this point another researcher, Noam Shazeer (now back at Google and working on their Gemini model), got involved and worked his magic to turn the idea into a realization - the transformer architecture - whose language modelling performance was indeed an improvement. Actually, there seems to have been a bit of a "throw the kitchen sink" at it approach, as well as Shazeer's insight as to what would work, so there was then an ablation process to identify and strip away all unecessary parts of this new architecture to essentially give the transformer as we now know it.
So this is the history and reason/motivation behind the transformer architecture (the basis of all of today's LLMs), but the prediction performance and emergent intelligence of large models built using this architecture seems to have been quite a surprise. It's interesting to go back and read the early GPT-1, GPT-2 and GPT-3 papers (ChatGPT was intitally based on GPT-3.5) and see the increasing realization of how capable the architecture was.
I think there are a couple of major reasons why older architectures didn't work as well as the transformer.
1) The training efficiency of the transformer, it's primary motivation, has allowed it to be scaled up to enormous size, and a lot of the emergent behavior only becomes apparent at scale.
2) I think the details of the transformer architecture - interaction of key-based attention with hierarchical processing, etc, somewhat accidentally created an architecture capable of much more powerful learning than it's creators had anticipated. One of the most powerful mechanism in the way trained transformers operate is "induction heads" whereby the attention mechanism of two adjacent layers of the transformer learn to co-operate to implement a very powerful analogical copying operation that is the basis of much of what they do. These induction heads are an emergent mechanism - the result of training the transformer rather than something directly built into the architecture.
We had a talk about those physics AIs using those maths AIs to design hard mathematical models to fit fundamental physics data.
"|" "View" "ing" "Single"
Just looking at the text being tokenized in the linked article, it looked like (to me) that the text was: "I View", but the "I" is actually a pipe "|".
From Step 3 in the link that @miletus posted in the Hacker News comment: https://x.com/0xmetaschool/status/1888873667624661455 the text that is being tokenized is:
|Viewing Single (Post From) . . .
The capitals used (View, Single) also makes more sense when seeing this part of the sentence.
To be clear, neither possesses any magical "woo" outside of physics that gives one or the other some secret magical properties - but these are not arbitrary meaningless distinctions in the way they are often discussed.