Like counting the number of R's in strawberry, many of these are character-counting or character manipulation problems which tokenization is not well-suited for.
I'm sure an engineer could come up with a clever way to train for this, but that seems like optimizing for the wrong thing.
IMO these questions go in the wrong direction. Character permutation is a problem for "Software 1.0", not LLMs. Just as you wouldn't use an LLM to multiply 2 large numbers, you'd use a calculator.