Assuming that HumanEval is a good benchmark (it's not) and assuming you can naively scale under fat-tails (you can't) then according to OpenAI's own gpt4 report (someone did the math on hn/reddit where they reproduced the curve but i can't find the link) at 100x the training cost of gpt4 you will still have a 15% error rate on medium difficulty tasks.