What high quality data sources are not already tapped?
Where does the next 1000x flops come from?
What high quality data sources are not already tapped?
Where does the next 1000x flops come from?
Stick a microphone and camera outside on a robot and you can get unlimited data of perfect quality (because it by definition is the real world, not synthetic). Maybe the "AGI needs to be embodied" people will be right, because that's the only way to get enough coherent multimodal data to do things like long-range planning, navigation, game-playing, and visual tasks.
Some people don't seem to realize how critical the "eval" function is for machine learning.
Raw data is not much more useful than noise for the current recipes of model training.
Human produced data on the internet (text, images, etc.) is highly structured and the eval function can easily be built.
Chess or Go has rules and the eval function is more or less derived or discovered from them.
But the real world?
For driving you can more or less build a computer vision system able to follow a road in a week, because the eval function is so simple. But for all the complex parts, the eval function is basically one bit (you crashed/not crashed) that you have to sip very slowly, and it very inefficient to train such a complex system with such a minimal reward even in simulations.
I don't see how this is any less structured than the CLM objective of LLMs, there's a bunch of rich information there.
There is at least one missing piece to the puzzle, and some say 5-6 more breakthrough are necessary.
It is not like crashed/not crashed is the only possible eval function.
It can be easily much more nuanced than that. The driving system should be able to predict how everyone will move next is a good sub-goal. Checking if you were in the positon of an other driver, seeing what they see would our code be driving the same way as them is also a good sub goal. (Obviously total alignment here is neither possible nor is it desireable.)
Other evaluation is to check if you forced anyone to change speed/swerve to avoid you. And then you can have synthetic scenairos for every time you approached a lane which had priority over you. You can add conflicting vehicles approaching (with different timings and speeds) and see if own vehicle notices and handles them correctly. (And “handles them correctly” is not a binary crashed/not crashed either, you can check if the vehicle inconvenienced the simulated vehicle.)
Be careful with mistaking data for information.
You are getting a digital (maybe lossy compressed) samples of photons and sound waves. It is not unlimited, a camera pointed at a building at night is going to have very little new information from second to second. A microphone outside is going to have very little new information second to second unless something audible is happening close by.
You can max out your storage capacity by adding twenty ML high megapixel cameras recording frames as tiff file but gain little new useful information for every camera you add.
> Where does the next 1000x flops come from? Even with Moore's law dead, we can easily build 1,000x more computers. And for arguments about lack of power - we have sun.
Even still, we need evolutions in model architecture to get to the next level. Data is not enough.
LLMs can't do jack shit with ciphertext (sans key).