Where do they benchmark code candidates? This is hard, noisy and expensive (I see example reward for code size).
I am not an expert, but bayesian learning maybe more appropriate for such an expensive-sampling environment?
I am not an expert, but bayesian learning maybe more appropriate for such an expensive-sampling environment?
This works by changing around the order / interleaving of various LLVM optimization phases, so the learning process does not require knowledge of program timing or correctness.