Very strong statement on the title, given the following limitation:
> Generation tasks. Method applies to classification only. Preliminary decoder experiments show perplexity increases.
> Generation tasks. Method applies to classification only. Preliminary decoder experiments show perplexity increases.
The distillation of a student that predicts "anchor layers" and then acts as a backbone for classification is perfectly cool on its own; no need to stretch the title/abstract so much.
The title reflects the strongest verified result in the domain the method currently supports, not a universal claim across all modalities. In other words, the compression result is real, but it shouldn't be interpreted as applying to generative decoding... yet.