Self-hosting keeps your private data out of AI models
blog.zulip.com
blog.zulip.com
I don't know who else here has ever actually done the kind of work that we're talking about, but I have, for contract clients, over the span of most of a decade. Between G Suite and Office 365, those jobs dried up fast about a decade ago, and I wasn't sorry; the thing about selfhosted deployments of this type is that no matter how standard you try to make the platform, clients' requirements persist in varying, so each deployment ends up a somewhat different artifact whose differences, because those get less investment of effort than the common aspects, are always what cause the most problems and cost the most money in maintenance. No one was sorry to see that expense dwindle with the rise of SaaS officeware platforms; even for me, it wasn't a loss, because it freed up my time for more interesting and lucrative engagements.
Of course I'm sure that, for those of you saying this is easy, your experience is different from mine. It would be really interesting to hear more about how you've been able to make selfhosted deployments work so well for folks who aren't technical. That's a really significant achievement! Teams like Sandstorm have spent years at it and come away with nothing of import to show, and I'm sure I'm not alone in wanting to know more about how folks here have overcome the same challenges.
[Or used to, around 2006 or so - I'm not sure if they still claim managing public cloud resources is means less work.]
It would be nice if the tooling really was simple enough that wasn't so, but despite considerable effort to make it so from Sandstorm among others over quite a long time, as far as I know no one's had notable success in achieving that. Hence the rise and continued prominence of SaaS providers, who abstract all of that effort behind an SLA and a monthly fee.
I agree with the criticism of SaaS platform behavior at the head of this thread. What I think the commenters on that side of the discussion are missing or maybe ignoring, is that much of the value a SaaS offering provides is in not having to (pay someone to) administer that stuff yourself - and that, considered separately from the question of how providers behave, "all you have to think about is 'pay us and use our stuff'" is in the general case a strong value proposition.
There are new tools, apps and solutions every year (easy vm handling, kinda easy vpn, projects like Wireguard, Immich) but overall there are huge things missing to make selfhosting a thing for common people.
These documents are new in the last few days:
https://slack.com/blog/news/how-we-built-slack-ai-to-be-secu...
https://slack.com/intl/en-gb/blog/news/how-slack-protects-yo...
I think these updates are really good - Slack's previous messaging around this (especially the way they unclearly conflated older machine learning models with new policies for generative AI) was confusing and it wasn't surprising it caused a widespread panic.
It's now very clear what Slack were trying to communicate: they have older ML models for features like channel recommendations which work how you would expect such models to work. They have a separate "Slack AI" addon uou can buy that adds RAG features powered by a foundation model that is never further trained on user data.
I expect nobody will care. Once someone has decided that a company might "train AI" on private data you've already lost that person's trust. It's not clear to me if any company has figured out how they can overcome one of these AI training panics at this point.
I wrote a bit about this back in December when it happened to Dropbox - there is an AI trust crisis at the moment: https://simonwillison.net/2023/Dec/14/ai-trust-crisis/
I think it goes beyond a single company, or rather a single incidence of this panic. You're looking at each time this happens as an independent coin flip instead of a series of dominoes that trigger a reaction in multiple directions.
What I mean by that is there's a counterculture sentiment building based off the idea that people have seen this same pattern enough times at this point that they're distrustful of large scale systems by default. It's happening with government institutions, politics, economics, and individual industries like gaming and streaming.
To that end the "panic" is not just a reaction to Slack's (perceived) actions, but an expectation that Slack will be yet another domino in that line of companies that have done the same. It's also difficult to prove a negative (that Slack isn't using private data for training purposes even if they say they're not) so the messaging is up against a very solid wall.
The result here is that public announcements and messaging related to data are under heavy scrutiny, and the media is incentivized to try and make their reporting go viral (ironically for the ad revenue) at the expense of actual journalistic reporting.
I'm not sure what the solution to this problem is, or if there even is one, but promoting self hosting seems like an indicator that the default assumption is that data collected will be abused in some way. Honestly based on the last couple of years it's not an unreasonable assumption either.
The crux of the matter is whether you can trust a big tech company to do what they claim they will. They all think AI is worth infinite dollars. In that world, without some very clear, painful, straightforward, contractual penalty ... well we've seen that the tech giant plan is that rules are meant to constrain your competitors' behavior, not yours.
If they wrote "If any of your data is discovered to have been in an AI training model, Slack owes you 10x your lifetime payments to Slack, and any involved whistleblowers get 1% of the total paid" in their terms of service, which means if Slack screws this up, the company is immediately bankrupt, that might prove effective. But a promise in a "privacy principles" policy that doesn't appear to actually be incorporated into the core ToS does not have a lot of teeth.
I think you're right: the most convincing version of this would be actual legalese.
I wouldn't be surprised if Slack have this in the contracts they sign with their larger customers, but I don't think those are publicly available.
Now seems people's chats and posts are being used to train AI. I wonder how long before Cell Phone Providers start using Text Messages to train AI (or sell to AI people).
I've seen several folks have this opinion... and it just doesn't make any sense?
SMS for 2FA is problematic for many reasons, but Github offers TOTP & U2F/passkey 2FA, which have zero privacy implications.
I don't understand why would something like that be a reason to avoid github (or any service) in itself?
In this era, 2FA is important security measurement and it is not like they enforce only SMS as the only 2FA form.
Second, it's not an important security measurement for me. I didn't go into FOSS to be part of someone's supply chain. I do it as a way to share my knowledge of how one might solve a problem. If you want use my code, then inspect it to make sure it does what you want, or pay me for commercial support. Neither require 2FA.
Might someone take over the account? Sure, I suppose. But I'm not into "community building" or GitHub's gamification, and my primary repos are all local, so if that happens and GitHub's support didn't help, I could start a new account. Again, don't depend on me for your supply chain without a commercial support agreement.
When Microsoft switched GitHub to require 2FA I concluded it was because they wanted to assure their corporate and government clients that it was "safe" for them. Those profits subsidize Microsoft's free hosting plans, so my presence there was helping contribute to Microsoft's already excessive market power.
Third, the change was driven from on high, with no chance for me to decide what was appropriate for my projects. I concluded Microsoft was so powerful they could make such paternalistic changes because they knew network effect was on their side that they could have little concern about the small number of people leaving or getting upset.
Fourth, my FOSS projects on GitHub were labors of love that were a net negative on my income. I was not going to spend any money on new hardware or waste my time figuring out how to get things working under a new system when I was already hosting most of my work on Sourcehut, which is much more aligned to my ethical and moral views.
I still don't know how many security keys I'm supposed to have (how often should I expect to lose one? should I store the backups off-site at a friend's place?), or how often am I supposed to test they work? And then I hear about issues about lock-in and how attestation requirements might prevent FOSS solutions ad prevent people from backing up one's own security keys, and issues with resident vs. non-resident keys, and being able to register multiple keys. It's all learnable, but I simply don't care enough.
And I don't see why I should care about all this when the paying customers of my software have all been fine with only a tar.gz, license agreement, and support contract.
(I had 2FA set up a long time ago, so didn't notice the policy change)
IT Sec and Compliance must read the T&Cs and make better vendor selection choices.
Bashing a competitor in a blog post does not.
So in my opinion, it just doesn’t matter if you are using self-hosted AI, the weakest link in your chain for keeping your data private is the very OS’s that you’ll be interacting with said self-hosted AI.
And with all the manufactured fear mongering going on around AI, that data will -already- be deliciously irresistible for prism-participating, lovable, trustable companies like Microsoft.
Sorry to burst some pretty bubbles for the lovely naive people.
:). What a clever way to say that even though we don't do it today, we cannot guarantee that we will never do it on our cloud service. At least they are honest I guess.
I think you're missing a big part of the point of the post: Which is that if you self-hosting, nobody can train models on your data, if you're going to use a cloud service, you should use one where you can move data to self-hosting, and where you can trust the vendor.
We do basically guarantee that you won't have your organization's data included in AI training in Zulip Cloud without consent. But yes, we're not ruling out the possibility of some sort of opt-in feature that might be useful in Zulip Cloud.
I'm humble enough to not pretend I know what will be possible/expected in terms of AI technology in 5-10 years, but one could easily imagine some sort of tool trained on web-public channel data in open Zulip communities being a thing that could make sense if done with appropriate consent. If such a thing were desired, I don't think it most of the concerns related to the slack controversy would apply.
Zulip's weasel wording indicates they are nerfing a great (maybe their best) opportunity to stand out from the herd.
How about this for a mind-blowing concept (/s) ... If a web-public channel wants to add some sort of useful feature based on a technology trained on their data then let the owners/administrators of that channel flip that switch on. Zulip should have no involvement in that decision, period.