HuggingFace: Security.txt
huggingface.co
huggingface.co
The model also lacks the machinery. No training loop, no gradient descent, nothing to write to.
And a model only sees its own sampled tokens, not the distribution behind them, which are possibly filtered or post-processed. Distillation from that works but is less sample-efficient than soft-label distillation.
Now I want to go find that conversation in my history and see if it can tell me about weights, and how that contributed.