If you're like me you'd probably start with some kind of "grid search" by giving it evenly spaced parameters. Then you would evaluate it at each one and do some kind of gradient descent / ascent type algorithm. This will work, but in a high dimensional space (ie. a function with many arguments) you have no chance of covering even a tiny fraction of the parameter space. The key is, you don't want to waste time evaluating a set of parameters if it doesn't tell you anything new. ie. if you get back a similar answer to what you already had. This is what the GP means (I think) by trying to avoid "dependence between adjacent draws".
I could be wrong about all this, I'm trying to learn it too.