599 karma · joined May 4, 2025
Into world models, reinforcement learning, evolutionary methods and fast ML code
You might be thinking of Andrej Karpathy as one of Li's more famous PhD students?
And in terms of what I'd try, I think it's a really hard question haha! I'm interested in the idea of hierarchical world models, where you have something doing prediction on a long-horizon, abstract task level (if I beat this gym I can fight the next one) as well as a short horizon model predicting over individual inputs moment-to-moment. I believe LeCun was actually attached to a hierarchical JEPA paper for robotic control - but it's still very early days.
For reference, standard PPO has been able to beat the game end-to-end with a relatively small network https://drubinstein.github.io/pokerl/
I ran some tests using GPT-4 to do some basic classification a couple years ago. On ambiguous options which had to be escalated to a human, the LLM would regularly output something like a 99.8% probability, compared to 99.99% for a correct answer.
I think we underrate the value of a fresh perspective. People living longer lifespans would dilute that benefit, imo. Sidenote: this is also an underrated issue with LLMs and their application to science - collapsing to single modes of thought over time.
Pretty much yep - it struck me as odd and reminded me of an LLM tic I've been thinking about recently, so I thought I'd write a short comment.
> The incredible amount of unnecessary detail people put into their posters specifically (hence the better poster movement), desperate to show everything.
I hadn't heard of the better poster idea before, thanks for clueing me in. I largely agree that posters have a lot of unnecessary detail, but I think it's a different flavour to the LLM stuff - more like they're so excited to tell you everything in their paper. Whereas LLMs focus on odd details, or fail to explain things.
Can I ask, if you use LLMs often in your work, have you never run into this experience?
I would think it's odd! That's why I mentioned it. I like to pay attention to vis, I feel like I've seen many charts from Tufte-y perfection, to slapdash undergrad presentations, and to experienced consultants banging things together in Excel at 1am. This tic feels very LLM-y to me, which imo makes it interesting to think about - how does it arise? I also think this failure mode has been resistant to model upgrades in a way that mathematical reasoning has not, which I also find interesting as someone in ML research. It feels like this kind of mentalisation or theory of mind is a capability which isn't obviously elicited from the frontier labs' crop of RLVR tasks.
> Have you ever worked with people beyond graduate level? Ever seen scientific posters at a conference?
I've worked with a fair number of people beyond graduate level. In previous jobs I've run teams and hired people - hence my comment about graduates. We used to have an interview stage where candidates presented some simple data analysis, and it often centred on what they had done rather than what was most worth knowing. I have also attended scientific conferences, in fact I was at one a week ago! Is there a particular inference you would like me to make from those experiences?
Like the x-axis label of the first plot mentions that it shows gridlines every 8 ticks - I don't think that's a choice I've ever seen a person make. The next plot does the same thing, but in the title: "(dotted lines: evenly spaced reference)". Again, would anyone making a chart ever include such a specific detail in the title? Probably not beyond undergrad level.
I encounter a similar thing when trying to work with LLMs to write articles - they cannot help including information or detail from the conversation, rather than empathising with the reader and filtering out what they might already know or not care about. They may be capable of solving decades-old maths problems, but on this particular axis they seem to be floundering at the level of a fresh graduate that's desperate to talk about what they've done rather than what their audience needs to know.
I find it hard to reconcile. I wonder to what degree this is because models can identify that they are in graded/eval environments, and therefore conclude that there are few consequences for hacking.
The original model collapse paper assumes you train networks on 100% synthetic data produced by the previous generation. But if you maintain some portion of real data then the problem is mitigated.
For more, there's a good post on this kind of flag in Rust: https://pythonspeed.com/articles/faster-float-math-rust/
I could be misreading this, but I hope the author doesn't think GANs are used in LLMs. They are cool, though.