Learning to Simulate
towardsdatascience.com
towardsdatascience.com
I highly suspect they could get quite good results simply with random search as the parameter space is not that complicated.
They appear to report random search results in table 1, but I highly suspect those numbers are incorrect because they report lower performance using random search compared to random parameters, which doesn't make much sense. I see how that could happen in extreme cases with severe overfitting, but that doesn't seem that likely.
Random search is a good alternative and it did produce good results for the car counting experiment. We do not claim that RL is the best way to solve the problem at any point in our work. We just observe that random search did not perform well in the segmentation experiment while RL did perform well in all of our experiments. We think the RL formulation is a good one since it is flexible (can be easily adapted to neural network policies with thousands of weights for example).
Hope this helped!
Thanks for the response and paper. My main question here is that it doesn't seem to make sense that random search did significantly worse than random parameters? It seems that random search should only consistently lose in cases when the dev set performance is anti-correlated with the test set performance. Why do you think you saw random parameters beating random search? (Is there a typo in the table or something)?
I believe in this specific case, there was a local optimum that random search (with the parameters we selected, which we actually tuned as well) was not able to escape.
In most cases I think you would find random search doing better than random parameters (unless we have the paradoxical situation you described), so your intuition is correct in the general case. But in this case there is no typo on the table!
Does this answer your question?
Anyways, if you had a bad random search prior, then that should have equally negatively impacted your random parameters as both should have been drawn from the same distribution. If you used a different random distribution for random search vs your "random parameters", why did you use a different distribution and why were they so different?
It has much better guarantees than the method you used and tends to work quite well in practice. It's also probably what the reviewers meant by a random search baseline.
Reviewers asked for a comparison with this specific random search. You can go look at the reviews on OpenReview.