Fwiw, nobody has ever suggested to me that I employ token compression in my daily workflow. I don't pay full attention in all the AI workflow demos I'm supposed to attend, but I don't recall that even being discussed. Is this an Nvidia blog or tweet you're referencing? I'm actually interested to see what they have to say.
Please don’t tell me you’re writing RTL
I'm not, I work higher level products. I've talked to a few people who do but I don't recall if they have different standards.