Reinforcement learning does work for this, but it's brittle. You almost always end up with a strategy that exploits inaccuracies in the physics simulation – from the perspective of the RL algorithm, useful quirks of the laws of physics – and so doesn't transfer to reality very well, if at all. Its behaviour outside the conditions observed in training is not guaranteed (or even expected) to be sensible, and even its behaviour within the training conditions is often hard to characterise.
An algorithm designed by people who know what they're doing is usually better. More effort, yes, but rockets are a lot of effort! We can afford to pay the cost for a reliable landing system.