If they're trained on the output of a larger model...
Would they inherit "thinking" that they know more than they do?
It'd be interesting to see if distillation increases hallucinations for specific topics the larger LLM is confident in
It'd be interesting to see if distillation increases hallucinations for specific topics the larger LLM is confident in
No comments yet.