Minimaxir thanks for comment.
Regarding "accuracy" I agree that this word usage is unfortunate in case of R2. A "goodness" would be better description of R2.
Regarding the "one-hot" encoiding of inputs, I would say - it depends. I know it is a common practice to do "one hot" encoding in Machine Learning. There are cases when it really helps the model to converge (for example, sometimes Neural Networks would require such curation of input).
However, model that I have used in my post is Decision Tree based. And when you think for a moment about it, Decision Tree based model should handle numerical input just fine. Usually, Decision Tree based models do not even require standardization/normlization of input and work correctly without it.