Mathematical Foundations of Reinforcement Learning
github.com
github.com
I'm absolutely not versed in RL, but I wanted to understand GRPO, the RL algorithm behind Deepseek's latest model.
I started from a very simple LLM, inspired from Andrej Karpathy's "GPT from scratch" video (https://www.youtube.com/watch?v=kCc8FmEb1nY). Then, I added onto that the GRPO algorithm, which in itself is very simple.
I made a GitHub repo if you want to try it out : https://github.com/Al-th/grpo_experiment
Very curious about RL for LLMs for example (using data from real use).
Neither cover LLMs. I don't follow the literature closely so I can only suggest you read papers: https://github.com/WindyLab/LLM-RL-Papers
Not trying to catch you, genuine interest.
Im expecting to see a lot of hyper growth startups using RL to solve a realworld problem in engineering / logistics / medicine
LLMs currently attract all the hype for good reasons, but Im surprised VCs dont seem to be looking at RL companies specifically.
The period from ~2012-2019 of AI research had deepmind (who was the undisputed leader in money and talent) go all in on RL to solve problems and while they did do lots of interesting and useful work, there wasn't anything quite so extraordinary / revolutionary in massively accelerating the field or some sort of crazy breakthrough.
Their over-focus on RL instead of transformers/llms is what allowed OpenAI to surprise everyone and overtake deepmind.
Yes, RL is a useful tool, but outside the context of training LLMs for reasoning there isn't really any breakthrough that makes it more than an interesting tool for certain situations.
I love when people on hn make market predictions based on how revolutionary they think something is. I guess startup people thank they're also VC people.
FYI Sutton's book came out in 1999; none of this is revolutionary anymore and yet I don't see any "hyper growth". The reason is exactly because while you can train these models to play super Mario, you cannot use them to solve real world problems.
https://www.google.com/books/edition/Reinforcement_Learning/...
Perhaps thats because it takes a while for the ideas to get polished/weeded and diffuse into the engineer zeitgeist .. or it could be that compute / GPUs are now powerful enough to run at the scale needed.
re : "RL cannot be used to solve real world problems" .. well, I would argue that these are useful real-world problems :
- predict protein folding structure from DNA sequence
- stabilizing high temperature fusion plasma
- improving weather forecasting efficiency
- improve DeepSeek's recent LLM model
Im currently using RL techniques to find 3D geometry - pipes, beams, walls - in pointclouds.
It is of practical benefit, as a lot of this is done manually, ballpark $5Bn/yrBut I concede I cannot point to a plethora of small startups using RL for these real-world problems .. yet.
This is a prediction, and I could be wrong in many ways - not least that LLMs digest RLs in full and learn to express their logical reasoning, approaching AGI, and use RLs internally, and so subsume and automate the use of RL.
Are VCs better at predicting the future.. I guess that is their job, and they have money on the line... but I think even they would admit they need a large portfolio to capture the unicorns.
VCs probably get a less detailed tech view than founders, but the large number of pitches they review should give them a noisy but wider overview of the whole bleeding edge of innovation.
I think startup founders are in the same future prediction business .. and arguably have more skin in the game.
Predictions would be pretty useless if they weren't somewhat controversial - a prediction we all agree on doesn't say much. Come back and chastize me if we dont see more RL startups in 12 months time !
1999 is 26 years ago but ya sure this is the year they finally take off.
> Perhaps thats because it takes a while for the ideas to get polished/weeded and diffuse into the engineer zeitgeist .. or it could be that compute / GPUs are now powerful enough to run at the scale needed.
Or perhaps it could be that you're wrong and they're useless? Nah that couldn't be it.
Doesn't waymo and other self-driving systems use reinforcement learning? I thought it was used in robotics as well (i.e., bipedal, quadrupedal movement).
however multi-armed bandit algorithms are highly useful in practice. these are a special case of RL (RL with one state, essentially).
there are even some extensions of applied bandit algorithms to "true RL", e.g. for recommender systems that want to consider history.
this is the place to look for real-world applications of RL.
also RL uses importance-sampling estimators of the gradient. these sometimes show up in other applications though not framed as "RL".
This is so funny to me, I see it often and I'm always like "yea, right, some knowledge"... these statements always need to be taken with a grain of salt and an understanding that math nerds wrote them. Average programmers with average math skills (like me) beware ;)
And then you work on .net/java/sql/server crap for a decade and you forget even the little math you used to know :D
- Do you understand the material?
- Can you utilize your understanding to build successful models/algorithms?
If the answer is yes to both, do some projects, put them on your github, and update your resume. You might need to take a job at a lower position first, but you can jump from there. But I want to make sure that the answer is "yes" to both and note that it is easy to think you understand something without actually understanding it. Importantly we must recognize that everyone has a different level of sufficient knowledge where they are comfortable saying that they "understand" a topic. One person might say they don't and be more knowledgeable than someone that says they do. But demonstration of the knowledge levels is at least a decent proxy for determining this.A way I like to gauge someone's understandings of things is by getting them to explain the limitations. This is often less explicitly stated in learning and a deeper understanding is acquired through experience and most importantly, reflection on that experience. This is often an underutilized tactic but it is very effective. If you can't do this, then the good news is that starting now will only accelerate your understanding :)
Understanding the limitations is a complicated thing in tech. You can finnangle most systems into doing mostly anything, as inefficient as that may prove to be.
The question then becomes up to what point is it "a reasonably better than most others" solution. And that's a question of an understanding of a field, not a space in the field.
> is a complicated thing in tech
That's the point. Understanding complex things is what experts are supposed to do. > You can finnangle most systems into doing mostly anything
"most" is doing a lot of heavy lifting here and I think the point you're making isn't discrediting my point. Sure you can hamfist a lot of things into working but an expert should know when to use better tools. Being able to identify what would end up as a very hacky solution from one paradigm but could be efficient and/or elegant in another is what an expert should be able to identify. Essentially, are they able to reduce technical debt even before that debt is taken on? > an understanding of a field, not a space in the field.
Would you mind clarifying the difference? I agree these are different things but I'm not sure why understanding the limitations would imply not having narrower domain knowledge. Sure, in ML knowing the advantages of convolutions over transformers and vise versa is good. But if you're working on LLMs, ViTs, or anything else it is still good to know what the limitations of transformer models are, and specifically what attention can and cannot do. We should be able to get more and more narrow too. An expert will be able to understand the nuances of specific evaluation methods: metrics, measures, datasets, and other forms of analysis. Being able to discuss nuance and detail is how you determine if someone has expertise or not. IME it tends to be pretty easy to identify experts (even in other fields) due to their ability and frequency of discussing nuances.Having done research in RL, a big problem with incremental research was to reproduce comparative works, and to validate my own contributions. A simple library like this, with built-in tools for visualization and a gridworld sandbox where I can validate just by observation, is very helpful!
https://bfskinner.org/wp-content/uploads/2020/11/978_0_99645...