As the customer base becomes more and more corporate (which it will), they end up with disproportionately more customers whose experiences cannot be used to train the model to make it better for those customers.
Either way, corporate customers cannot leach off the training from consumers handing over their personal data forever; there aren't enough specialists in that training set to improve the models with no loss of corporate trust.
Betrayal of their trust is inevitable.
At some point, where does the training advantage for specialist LLMs come from, if not progressively encroaching on customer data for the benefit of equivalent customers?
I’m not making any accusations, but we should not underestimate their tolerance for legal and financial risk.
It may be a little paranoid to insist on self hosting based on that, but I’m not so sure that it’s crazy.
Which they did do, but scale is relatively miniscule to the full dataset.