I agree completely a more complex benchmark should be done with a complete cross-validation.
Just for future reference I did ran the fitting a few times founding very(+-2%) similar results. Also Random Forests do an average so probably not much to improve on that particular algorithm.