Privacy is an afterthought. Here's how devs can easily make it better.
stackoverflow.blog
stackoverflow.blog
- dont collect data you do not absolutely need to service the user
- do not use third party libs or services, where you do not understand how they handle the data you submit to it
Am I missing something here?
This is of course a legal document and the implementation may do something else.
So the question to answer is how can we ensure an interoperable contract for data between systems/services - that requires an ontology for privacy that makes enforcement easy(er).
It is possible to make privacy definitions a declarative and low effort part of development for engineers - then code becomes the enforcing layer instead of legal agreements.
While the latter part is the prime responsibility of the developer team. You need a culture of skepticism towards 3rd party access to you (customers, users, company) data.
I'll move if my customers move too.
If nobody wants to move first, we have to all vote to move at the same time.
https://slatestarcodex.com/2014/07/30/meditations-on-moloch/
> There’s a passage in the Principia Discordia where Malaclypse complains to the Goddess about the evils of human society. “Everyone is hurting each other, the planet is rampant with injustices, whole societies plunder groups of their own people, mothers imprison sons, children perish while brothers war.”
> The Goddess answers: “What is the matter with that, if it’s what you want to do?”
> Malaclypse: “But nobody wants it! Everybody hates it!”
> Goddess: “Oh. Well, then stop.”
I'm conscious that it's easy to say "do _____ better" and insert a soap box like privacy or security; but if you don't provide developers with tools, whether that's libraries, IDE plugins, linters or code analysis tools to make that task easier, it's almost impossible - can't ask developers to be privacy experts, but we can help them to make that expertise readily available.
Like the example you took - provisioning cloud infrastructure is a heck of a lot easier now that it was a decade ago because of a bunch of orchestration and infra as code tools that ease the burden/knowledge gap for developers. We've got to have the same for anything we care about - in my case, that's privacy.
If I can offer my two cents: if we do better at annotating data at source of ingress (whether that's data provided by users, inferred about them, etc.) such that we can better describe the data we hold, why we have that dataset and for what limited uses, we can then enforce those conditions on models - right now that additional context just doesn't exist so privacy type enforcement on ML becomes arbitrary and subjective based on a teams needs. We can do so much better if just describe what, where, how and why we've collected data in our systems - then enforcement is layered on top of that.
YMMV of course
Candidate: Will this job require me to violate anyone's privacy.
Hold on, how about this instead.
Candidate: Will this job require me to do anything illegal.
Candidate already knows the answers to these questions. She has no need to ask.
This "the boss made me do it" defense seems to be a recurring comment on HN, perhaps from those with a guilty conscience. But is it really persuasive. It is like asking the reader have empathy for a drug dealer selling fentanyl because "he really needs the money and no one else will hire him".
"I really just want to sell marijuana but the people higher up the chain decided we should sell opiates instead."
Otherwise how do we know a bridge is structurally sound enough to walk across? We have to trust in the checks/balances, systems and people that work on that infrastructure. The same should apply to any software that affects large numbers of people.
If a company is setting out to do ill, that's a societal failing we should all care about and hope to prevent, but it won't be solved just with technical measures.
Strong privacy is a kind of anachronistic issue in that the regulations came long after many systems were designed/built but also most of the common methods used to continue to design systems. So consistent data deletion across distributed system should be easy, but of course in truth as you know, it's a nightmare and often a brittle solution that needs to be updated as your software continues to change.
This specific example you've taken is a really good one of how painful privacy can be and how avoidable this issue should be for all devs, both software and data teams.
You're right that it's at every layer in any complex organization but I'm bullish on the belief that developers hold a lot of the solutions here and can make it happen even in a complex organizational hierarchy with competing interests. We first need better tools to make it easier to implement.
Exemplary punishment for any security exploit gone wild.
Management will start getting the required resources to make it happen accordingly.
For example, what if "management" is a one or two person startup?
Maybe punishment is not the answer, but rather liability insurance coverage requirements. Or treat it like workers compensation where a small tax funds an insurance pool. And make it so repeat offenders get charged an increasingly higher tax rate.
To take the restaurant example, a mom/pops restaurant may have less resource to bear for cleanliness and safety but if it consistently, knowingly persists in doing something that makes it's patrons ill - it is any less at fault than a chain of restaurants that does the same? The fine may be proportional to that organization - that's the goal with the GDPR's revenue % based fine format but it could/should go further for large companies that consistently fail.
When kitchen cleanliness, plumbing, food quality and preservation, cutlery, access for disabled people, ... becomes an afterthought, it is time to be shutdown by consumer protection government agency, usually they get one time warning though.
Or maybe not, depending on the country, but then expect what might be great food with interesting side effects.
However, I would argue that there are plenty of very small companies that also take advantage of that. i.e. very high growth, early stage companies with loads of vc backing that don't prioritize this because they're small and there's only two founders.
All I mean to say here is, again I'm not a regulator, however we do need a way to enforce against bad behavior.
In the 1950's no one wanted seatbelts, not car owners/public or auto manufacturers. Today no one would get in a car without seatbelts without thinking it was weird/crazy. Sometimes we have to enforce rules to drive change, otherwise bad behavior (particularly at large companies) goes unchecked.
If that's the case, we probably differ on POV a little so I'd love to know more about why you think this?
In my experience average people who don't work in tech (non devs) do not understand privacy or data use in a system - so it's hard for them to comprehend the potential impact. They're trusting people that work in tech (devs and others) to do the right thing and this is where the problem might be; we're being trusted to ensure we don't abuse or accidentally misuse a position of tremendous knowledge and power.
My parents certainly understand where/how their data is stored or used when they use their phone - aren't we then responsible for keeping those people who don't know safe?
TLDR; I agree with you, but I'm not a politician and can't effect change there, so I'll keep chasing a realistic solution that makes it easier for devs to do of their own accord as we (dev community) are pretty good at solving things when we turn our mind to it.
To operate, a phone company might need to know where you are calling from and who, a doctors office might need to know your medical and contact info, an isp might need to know your ip address, a dating website might need to know your ip address, your chat app might need to know your contact list, your gps might need to know your precise location at this specific time.
Do they need to know them for years though? And once all this info is aggregated, how personal is the information that can be learned?
Organizations have poor/no data retention policies so they accumulate information they no longer need for many years and on the other side they continue to build new data processing capabilities that might leverage that old data which lacks any context as to what it was gathered for or how it can (or cannot) be used.
On the point of what degree of identification can be learned (i.e. how personal is the information), I'm constantly reminding by folks in the ML field that you can discern a huge amount about an individual quite precisely from what might initially look like anonymized data - that's one of the exploits that most concern privacy specialists when they talk about differential privacy and pseudonymization - arguably we're not where we need to be yet but thankfully there are a number of teams working on solutions to this.
By this I mean the data analysis part; the retention and policy issue, that's still one most companies need to do better on.
Before product-market fit, you don't really know what data you need and what data you'll need. You can't go back in time to collect data, you need to do it now if there's a chance it will be useful later. In a pre-seed company (or even at seed), you don't have the resources to audit every package much less get SLAs in place. Most companies I know do the bare minimum for GDPR, and it's not a lack of care for the user, but that often it's not the best place to deploy resources when considering the survival of a company.
Most software starts out being assembled in flight with multiple changes in destination - and those habits stick.
I still meet so many who are laboring under the impression that it won't happen to them - it's just not the case so we all have to have a privacy and security first mindset.