ParentFull threadrockinghigh·For small models, tokenization can reach 1-10% of total inference time.View on HN