> their main trick for model improvement is distilling the SOTA models
Could you elaborate? How is this done and what does this mean?
Could you elaborate? How is this done and what does this mean?
I also haven’t seen any hard data on how much they do use distillation like techniques. They for sure used a bunch of synthetic generated data to get better at reasoning, something that is now commonplace.