I don't know how big "textfile.txt" is, but try it with a much larger body of text? Also you'll almost surely want a much better tokenization algorithm than that. Zipf's Law can only hold "at the limit" of an infinite stream of text, so as a consequence you get a better approximation with a larger sample.