That’s nuts and brings forward the idea that an AI is close to self improvement.
That’s nuts and brings forward the idea that an AI is close to self improvement.
Separatelt, but on a related note - do we need analog or quantum computing to "truly" scale?
You could have lots of comments and pass statements and unnecessary conditionals.
I don't think it matters that there is only one correct answer.
It matters that you can verify reliably if the answer is correct enough.
Not saying that it's impossible, but reinforcement learning of the form shown by DeepSeek was particularly well-established and robust. It's sort of ironic, actually...IIRC, OpenAI started in the world of this kind of reinforcement learning.
So why can’t we give LLMs other virtual environments to verify if the solution is correct? For example, a stock simulator, a physics simulator, a driving simulator, etc.
AlphaGo Zero is a version of DeepMind's Go software AlphaGo. AlphaGo's team published an article in Nature in October 2017 introducing AlphaGo Zero, a version created without using data from human games, and stronger than any previous version.[1] By playing games against itself, AlphaGo Zero: surpassed the strength of AlphaGo Lee in three days by winning 100 games to 0; reached the level of AlphaGo Master in 21 days; and exceeded all previous versions in 40 days.[2]
Training artificial intelligence (AI) without datasets derived from human experts has significant implications for the development of AI with superhuman skills, as expert data is "often expensive, unreliable, or simply unavailable."[3] Demis Hassabis, the co-founder and CEO of DeepMind, said that AlphaGo Zero was so powerful because it was "no longer constrained by the limits of human knowledge".[4] Furthermore, AlphaGo Zero performed better than standard deep reinforcement learning models (such as Deep Q-Network implementations[5]) due to its integration of Monte Carlo tree search. David Silver, one of the first authors of DeepMind's papers published in Nature on AlphaGo, said that it is possible to have generalized AI algorithms by removing the need to learn from humans.[6]
Google later developed AlphaZero, a generalized version of AlphaGo Zero that could play chess and Shōgi in addition to Go.[7] In December 2017, AlphaZero beat the 3-day version of AlphaGo Zero by winning 60 games to 40, and with 8 hours of training it outperformed AlphaGo Lee on an Elo scale. AlphaZero also defeated a top chess program (Stockfish) and a top Shōgi program (Elmo).[8][9]
To the extent that someone can figure out a way out of this epistemic box, I'm interested.