Docusign just admitted that they use customer data to train AI
twitter.com
twitter.com
There's no sure-fire way to automatically anonymize arbitrary customer data.
There's no sure-fire way to automatically anonymize arbitrary customer data.
There's no sure-fire way to automatically anonymize arbitrary customer data.
You can't anonymize it by hand. You can't anonymize it by machine. If you're taking blobs of customer data and dumping it into a machine, customer data can come out of the machine. Neither humans nor machines are error-free. If you believe DocuSign is going to do a perfect job, I've got a bridge to sell you.
Nothing in there is anywhere near impossible unless the system doesn't have the data.
If reliability is zero, so is the risk.
But understand what you are doing: you are adding noise to individual's data to make them less identifiable within a specified set of data. The amount of noise you need to add depends on the size of the set they are within.
The OP is also right: For arbitrary customer data (within arbitrary sets of data) you can't guarantee anonymization.
There are also multi-party computation systems, homomorphic encryption and federated learning approaches that all provide different approaches for parts of the private machine learning problem.
It's hard.
https://en.wikipedia.org/wiki/K-anonymity
k-anonymity is an attempt to solve the problem "Given person-specific field-structured data, produce a release of the data with scientific guarantees that the individuals who are the subjects of the data cannot be re-identified while the data remain practically useful."
One serious problem with that approach is that no matter how well you k-anonymise your data set, someone else's dataset may deanonymise yours.
I have a couple of external SSDs and USB flash drives start to flash their busy lights after letting them idle for a couple of minutes. These "bursts of activity" generally is proportional to the I/O work they have done before the idle period.
One of my flash drives did that GC/zeroing thing when it's ejected. After ejection and drive disconnect, its busy light flashed for some time depending on the work needs to be done. If you didn't hammer it when you used it, it did nothing, but if you did tons of work, it worked up to a minute or so after being ejected.
To my understanding, the idea behind the heat death of the Universe is that nothing interesting happens anymore, not that there is absolutely nothing.
Sorry, are you talking about the heat death of the universe or the last five Marvel movies?
2. The heat death of the universe does not mean one gigantic black hole. I’m just a hobbyist but my understanding of the theory is that black holes will continue to form, but through Hawking radiation, they eventually radiate out all their energy until it is all dispersed, ultimately leading to uniformity across the entire universe, max entropy, where “work” can no longer take place.
(It is an interesting question then whether information is actually destroyed through Hawking radiation?)
Thank you for saying that :)
If you're one of the lucky 10,000 that don't know the reference:
https://stackoverflow.com/questions/1732348/regex-match-open...
And if you're one of the lucky 10,000 that don't know the other reference:
Obviously there are potential hiccups. But a lot of good can be done with encryption and metadata.
Simply classify data into buckets for each value of each variable such that each unique combination of buckets for every variable contains at least N items for a suitably large number of N. It is then mathematically impossible to match an individual data point to any set of original data points less than N.
That's not what this data is and that's broadly not what LLMs are trained on.
and obviously partial information is very useful for exactly those with some other partial information.
this is why OSINT is very effective. it's possible to piece together enough constraints on these datasets that in the end the likelihood of a successful deanonymization is unexpectedly high.
What are the variables and what is the value of N when you're talking about "freeform text contracts containing PII"?
I know this is a bit of an absurd example, but I'm trying to understand the mechanics more than to provide a realistic counterexample.
The analogy here is: there are laws regarding confidentiality that probably were broken here.
While I agree that it is unacceptable to use customer data without consent (as suggested by OP's post), I disagree with the implicit assumption behind the comment that I responded to:
Namely the implicit assumption that human/biological intelligence/agents are somehow superior to artificial intelligence/agents in face of secrecy/confidentiality.
It boils down to the question whether it's possible or not to create algorithms that outperform humans in tasks involving secrecy and confidentiality.
While I can't think of any reason why that should not be possible in general, I agree that the current SOTA of generative LLMs is not sufficient for that.
Is throwing lots and lots of data and RLHF training on an LLM enough in order to make the probability of customer data leaks small enough to be acceptable?
I don't know. But I don't trust MBAs who salivate with dollar signs in their eyes to know either. And I fear that their lack of technical understanding will lead to bad decisions. And I fear those might lead to scandals that make Gemini's weird biases in image generation pale in comparison.
Yes, the user bears the cost when their confidential data is leaked and the company derives the economic benefit of mishandling it, which is why this keeps happening.
There was a young lady I had to reintroduce myself to every week or so. I think of her every so often.
I’m certain she doesn’t think of me.
(here's one for the page above: https://archive.is/e6K4K)
I find it interesting because it shows how different those two are.
Under feudal zero-sum world your question would make complete sense, but under a capitalist positive-sum world, it seems to me the answer is just "a lot of profit would be left on the table" (ie positive-sum mutual-benefit transactions), which explains why docusign would do it.
But also contract validation, like check if something missing or finding loopholes. If they are sneaky, they could sell that information to lawyers, who can definitely make use of it.
I'm not inherently pro-AI, but there are indeed valid and interesting use cases for such models, if the data can be anonymized correctly. And I'm not sure if it makes sense to single out DocuSign, because there are probably many companies which train their models on similar data.
Throw in any AI into a company and the vulture capitalists will load you up on cheap money
If you want your data to be private, encrypt it or better yet keep it in house.
This is no different from any "old school" ML fraud detection, customer journey analysis, etc., etc.
Has anyone really pondered the consequences or are we just repeating the concerns of others without thinking for ourselves again?
What are real actual consequences? Especially if I’m a small company? Except of course it repeating or regurgitating your content, which is a valid concern.
Unless I give my consent unhindered then my data is my data and cannot be used for any purpose at all unless I approve. This also means you cannot make blanket statements in Terms of use saying I grant you permission to use my data. I don't.
That is exactly it.
LLMs are nothing more than engines that predict the next token based on a mind-boggling large set of statistics of what tokens follow what. They're essentially Markov Chains on a massive dose of steroids.
If the training data includes one document that includes password rules like "The password shall be at least 12 characters long. The password shall include [...]", and then you have another document that includes a line of "The password is fdFjkkl1@!#", then if you ask the LLM to generate a document that includes password rules, there's a chance it may output "The password is fdFjkkl1@!#" in that document, because it knows that "is fdFjkkl1@!#" sometimes comes after "The password".
I guess we'll all find out if/when people use the model to get it to disclose other users information but I won't be holding my breath.
It's like when people found out that Ring/Alexa staff were viewing their recordings and got upset. Even if they make super sure not to commit any further violations with their access, the act of consuming and making use of my data is a violation in of itself. It violates my privacy, and my ownership of my data. Even if my data can't be retrieved from the product you use it to make, it still feels icky that you changed the terms of our deal without my permission, and took advantage of our relationship to make extra money off of stuff that is fundamentally mine, and I have no recourse.
If my landlord surveilled my apartment and sold the footage on PPV but only in China, the material consequences to my quality of life would be negligible. There's virtually no way that any of the people watching the footage would ever contact me or interact with me in any way, especially if they did it with all of their tenants and randomized the footage such that it's hard to track one specific person. But it would still feel super weird and bad, and it would still be a violation of our relationship.
If you have given contractual consent, we may use your data to train DocuSign's in-house, proprietary AI models. Your data will be anonymized and de-identified before model training."
This seems to be opt-in? Am I missing something?
That just means you clicked "I agree" at the end of 10 pages of legalese.
In their privacy policy (https://www.docusign.com/privacy) they say their lawful basis is "Consent (where required under applicable law)" which certainly sounds like if applicable law doesn't require it they would be use it without consent.
I'd suspect it can mean some some clause buried in contract document.
Also, it's some party's interpretation of that clause, not what some other party understood it to mean.
However, in this particular case, maybe most relevant is that I'd assume Docusign is facing a lot of corporate lawyers among its customers. So I'd guess the company can't get away with behaving like a stereotypical sketchy startup, towards its customers, and I'm sure Docusign has legal advisors who would tell them that if they ever needed to be told.
GDPR also has the concept of a "contractual necessity". A legal basis where data must be processed for the performance of a contrast. So say, you buy a t-shirt from an American online store and they need to process your shipping address. That's contractual necessity.
They seem to be mixing the two into "contractual consent", as if you are able to consent in a contract. Under GDPR, you can't!
In no other contexts can there be a more legible interpretation.
She had moved out of state, but either way I think they would've done it that way. Even skimming the EULA there was no way I was gonna esign through their service.
This confused the bank, and it was honestly a bit of a pain in the ass with some emailing back and forth. But they got me the forms and I signed analog, and it went through eventually.
Just sent them my I Told Ya So heh
By which I mean to point out the problem that the people signing the contract/document aren’t the ones choosing the signing service company.
> Thank you for the feedback. We are updating our AI FAQ to be more clear: We only collect data for training from customers who have given explicit contractual consent as part of their participation in AI Extension for CLM, AI Labs and select beta programs using AI. When we train models using customer data to improve the accuracy of our AI features, we only use data from customers who have given consent, and that is de-identified and anonymized before training occurs.
You can self-host
Beyond that "If you have given contractual consent, we may use your data to train DocuSign's in-house, proprietary AI models." -- I'm open to hearing that "contractual consents" actually means "literally all of you mwahahaha", but that's not borne out by the content.