Tokenizer training doesn't scale as well as model training, so general practice is to train on a subset of the full corpus.
Maybe if we’re talking terabytes it might not scale as well but so far in my experience training tokenizers has never been an issue. It’s training models that takes ages.