This is kind of just a measurement of how representative a language is in the distribution of the tokenizer training. You could have a single token equal to “public static void main”.
Seeing all the C languages and JavaScript at the bottom like this makes me wonder if it's not just that Curly brackets take a lot of tokens.
for (int index = 0; index < size; ++index)
instead of for index in 0...size
eats up a lot of tokens, especially in C where you also need this construct for iterating over arrays.`public` might have a token by itself, even though you can have `pub` occurring in other contexts, too.