> As of right now, we have no way of knowing in advance what the capabilities of current AI systems will be if we are able to scale them by 10x, 100x, 1000x, and more.
I don't think that's totally true, and anyways it depends on what kind of scaling you are talking about.
1) As far as training set (& corresponding model + compute) scaling goes - it seems we do know the answer since there are leaks from multiple sources that training set scaling performance gains are plateauing. No doubt you can keep generating more data for specialized verticals, or keep feeding video data for domain-specific gains, but for general text-based intelligence existing training sets ("the internet", probably plus many books) must have pretty decent coverage. Compare to a human: would a college graduate reading one more set of encyclopedias make them significantly smarter or more capable ?
2) The new type of scaling is not training set scaling, but instead run-time compute scaling, as done by models such as OpenAI's GPT-o1 and o3. What is being done here is basically adding something similar to tree search on top of the model's output. Roughly: for each of top 10 predicted tokens, predict top 10 continuation tokens, then for each of those predict top 10, etc - so for a depth 3 tree we've already generated - scaled compute/cost by - 1000 tokens (for depth 4 search it'd be 10,000 x compute/cost, etc). The system then evaluates each branch of the tree according to some metric and returns the best one. OpenAI have indicated linear performance gains for exponential compute/cost increase, which you could interpret as linear performance gains for each additional step of tree depth (3 tokens vs 4 tokens, etc).
Edit: Note that the unit of depth may be (probably is) "reasoning step" rather than single token, but OpenAI have not shared any details.
Now, we don't KNOW what would happen if type 2) compute/cost scaling was done by some HUGE factor, but it's the nature of exponentials that it can't be taken too far, even assuming there is aggressive pruning of non-promising branches. Regardless of the time/cost feasibility of taking this type of scaling too far, there's the question of what the benefit would be... Basically you are just trying to squeeze the best reasoning performance you can out of the model by evaluating many different combinatorial reasoning paths ... but ultimately limited by the constituent reasoning steps that were present in the training set. How well this works for a given type of reasoning/planning problem depends on how well a solution to that problem can be decomposed into steps that the model is capable of generating. For things well represented in the training set, where there is no "impedance mismatch" between different reasoning steps (e.g. in a uniform domain like math) it may work well, but in others may well result in "reasoning hallucination" where a predicted reasoning step is illogical/invalid. My guess would be that for problems where o3 already works well, there may well be limited additional gains if you are willing to spend 10x, 100x, 1000x more for deeper search. For problems where o3 doesn't provide much/any benefit, I'd guess that deeper search typically isn't going to help.