Can you clarify what you mean by 'non-differentiable outcomes'?
Differentiable functions are great because you can run gradient descent on them in a very optimized way. Example: if your objective is to have a very high value, search in the direction that has a mounting slope.
Though maybe I'm missing something 'cos it seems to me you can run gradient descend on non-differentiable functions. It just requires more evaluations.
Isn't that exactly the point though? If you don't have an analytical solution for the gradient of the loss (reward) wrt the parameters - yes - you could brute force a numerical solution but as the number of parameters grows that quickly becomes infeasible. Approaches such as RL and GA provide a more intelligent way to search the parameter space.