The fake claim here is compression. The results in the repo are likely real, but they're done by running the full transformer teacher model every time. This doesn't achieve anything novel.
If you look at the student-only scripts in the repo, those runs never load the teacher. That's the novel part.
What they achieved is to create tiny student models. Trained on specific set of input. Off the teacher model's output.
There is clearly novelty in the method and what it achieve. Whether what it achieve would cover many cases that's another question.