I'm working on a somewhat related product (except bringing this assistant capability to all apps on your computer, all browsers, and using mostly on-device ML...waitlist in my profile in case you're curious)
What we've discussed internally is having two modes for the cases where we do need network connection:
1) A turn-key, use-our-OpenAI/HuggingFace/whatever proxy that doesn't store anything, just adds our token and pays for it on your behalf
2) Bring-your-own key for each service
The fact is that most users who just want to use these kinds of productivity tools might not have their own OpenAI/Azure/etc account, so offering option 1 and even defaulting to it is right for most end-users.
I think XP1 is making the right call here with this default, though offering #2 would be nice!
(edit: added Github link to XP1)
I believe most people are uncomfortable with all data being sent all the time, but more okay with sending some data in the exact cases they choose, since they have control.
So I'd argue that the extension being open source helps a lot.
You're right that it doesn't guarantee anything for how the server is behaving in the case that you do invoke it though. For that we'd need either transparency into that source code and server operations, or, more likely, a strong privacy policy and maybe SOC2 or other certifications.
I believe the reason XP1 is subsidizing this right now is to grow their user base to attract investors, and as they develop their LLM platform, and then probably charge for business users down the road, but they don't seem to state that intention as clearly as they could.
I still don't think it's appropriate for them to be using responding to emails as an example in their docs, especially without a warning. If someone went around sharing my private conversations with another person without telling me, I'd lose trust in that person. They'll get away with it because some people don't see sharing data with a software company as the same thing, and it's tough to know when it's happening, but nevertheless, it's sketchy.
"Hi yunwal,
Regarding your concern about the privacy of Dust XP1, it's understandable to be cautious when it comes to sharing personal information. However, as mentioned in the Discord message shared by spolu, Dust XP1 only sends your requests to OpenAI and stores them for debugging purposes. They do not fetch or store anything else than what is required to process your requests. Additionally, Dust XP1 is open source, which means you can look at the code if needed. If you have any further questions, feel free to ask.
Best regards, XP1"
I'm not sure why this is necessary. Surely, they're just scraping the page and using it as context in a ChatGTP prompt. They don't have to proxy it through their servers to do that.
No, that's likely entirely made up by ChatGPT.
Hopefully that policy can assuage my concerns, since I'd love to use this! It looks similar to the Edge AI features Microsoft teased at the new-bing release, and I keep thinking how handy it would be to have that toolbox available while I'm browsing for work.
Privacy
Only the text content of the tabs you select and submit are sent through our servers to OpenAI's API. Cookies, tab list, or non-submitted tab content are never sent.
A bit less than what the model came up with.
Also, are stements made by an llm model about staments made on discord legally binding?
Are you genuinely curious, or are you asking because you're implying that the people who would use such an app are somehow not understanding something or not intelligent, or don't know something that you do? Like you have to prove your case too, as you're not immediately "right" in your statement. Sure there is some level of "risk" in doing this, but there is risk in a lot of things. It's like me asking people this:
"Why would anyone trust people with 4 weeks driving classes and a test with their lives on a road driving 80mph inside 2-tonne metal cages? Seems insane."
We've run the experiment for n decades with billions of subjects and the results indicate it is actually not that bad.
We have -not- really even discussed, much consider, the implications of having a central system with machine intelligence designed to extract features, patterns, emotions, assumed motives, ..., having access to the entire digital lives of societies.
Is it insane to repeat the same mistakes? Not sure but it's somewhere close in the neighborhood. We could run it by a k-nearest algorithm and see what that suggests as a better category than 'insane'.
Actually we've pretty much been watching that play out for years now, even if the technology hasn't been in its final form for the duration. Results have already been a wee bit society-destroying.
There is no business model here that doesn't include "We sell your most private and sensitive data to the highest bidder". Because it's free, they can't make money any other way. And while for other browser extensions/software, you can at least audit the requests being made by the extension, this thing is sending away all your data because it has to in order to work.