> We additionally observe other behaviors such as the model exfiltrating its weights when given an easy opportunity.
Is anyone familiar with how this occurs? Since the models can only output text, do they attempt to "connect" to some API and POST its weights?