Can we not just write tests and have some LLM try 10,000 different algorithms and profile the results?
Or is an LLM unlikely to find the optimal solution even with 10,000 random seeds?
Just asking. Optimizing x86 by hand isn't the easiest, because to think through it you start to have to try and fit all the registers in your mind and work through the combinations. Also you need to know how long each instruction combination will take; and some of these instructions have weird edge cases that take vastly longer or quicker to run that is hard for a human to take into account.