Pile-T5
blog.eleuther.ai
blog.eleuther.ai
You can try using GPT4's tokenizer on your own HTML inputs below [1] ... there's definitely room for improvement!
The modularity of the enc-dec approach is useful - you can insert additional models in between (e.g. A diffusion model), you can use different encoders for different modalities, etc
* llama tokenizer
* the pile dataset
* trained on 2T tokens (2 times of og T5)
* better, & especially better at coding