Another take on a similar idea from FAIR is a Large Concept Model: https://arxiv.org/pdf/2412.08821
In meta's byte-level model they made the tokens variable length based on how predictable the bytes were to a smaller model, allocating compute resources based on entropy.