Then again, what type of side-affects would that have on the quality of the products moving forward.
Then again, what type of side-affects would that have on the quality of the products moving forward.
Even if you follow best practices with access control, in the end you're always going to have a group of people who you need to trust with access to folks' personal data. Maybe the solution is better audit logging and even tighter access, but I'm not sure leaks of this nature are preventable.
This is absolutely not true. This number can and should be reduced to an absolute minimum number of people.
As in, define "tech companies" (name specific companies that do this and why you think that can be generalized to the entire industry) and define "careless" (it's a relative term so please say if it's less or more careless of other examples of organizations that manage similarly large amounts of data but that aren't "tech companies").
Because to me the opposite seems true. It's the non-tech companies (if you equate that to FAANG) that manage large amounts of data that tend to have a lot more data leaks (internal or external) than the tech companies.
These groups usually get access through a bastion that anonymizes data and logs access. I remember that as a SWE at Google, I could run aggregated & anonymized statistics across query logs, but some info (eg. IPs, user logins) had been scrubbed before any of my code could get access to it, and for things that were more personal (eg. your GMail login) you could only get access to your own account.
There's nothing you can do about SREs who have root access on the box or the SWEs who need to implement & maintain the bastion servers, but that's presumably a more restricted, vetted, and trusted group.
I would say it is much more likely that Google will accidentally lose the organizational ability to become root-in-prod, than it is that a person has done this thing without being noticed.
In short, insider risk cannot be mitigated with hiring practices. You need robust technical measures against insider risk.
https://www.businessinsider.com/google-engineer-stalked-teen...
Snowden's revelations (2013) were a major watershed. There'd been several measures taken since, based on what I heard on the outside, largely through discussions, mostly public, a few direct, with Google staff via G+.
Starting on, of all days, November 9th, 2016, I began regularly posting an image of Jewish shop windows shattered during Krystallnacht, asking whether Google were thinking of brownshirt-proofing their data. That generated responses including from G+'s architect (then in a role with user data safety & privicy), and the data security lead.
It wasn't until some time later that I realised I'd entirely accidentally picked the anniversary of the event for the post. Though the coincidence was useful.
My understanding was that numerous protections were in place by that time. I continue to have concerns.
This is so incredibly cringey. You're actively building the panopticon and yet you think of yourselves as righteous warriors for justice.
I don't mean this to be a personal attack, but yours is such a revealing comment about the mindset of people inside these surveillance behemoths.
(See also: this "pledge" http://neveragain.tech/ to not build registries for targeting citizens...signed by a bunch of people who work at companies whose entire business is targeting citizens with ads)
Sounds like you'd be surprised at what storage, and backup engineers have access to.
/s
OK, Google.
There was a video with some of the datacenter security measures, e.g. iris scanning (just to get yours in the DB required approvals from senior people). On the actual floor, to which very few have actual access, you need to badge both on your way in and out, individually. If you badge out without having badged in, the door won't open and an alarm will go off.
I think much like anti-hacking and anti-fraud efforts, publishing information about how they vet candidates would just make it easier for attackers to figure out how to game the system.
Why wouldn’t the information be public? It’s the same concept as crypto algos being published and peer-reviewed.
Of course many of the changes around that time were in response to the breach by the Chinese government, not only in response to embarrassing privacy incidents. Later improvements came about because of things Snowden published.
https://www.washingtonpost.com/world/national-security/nsa-i...
I suspect that at some stage this might come for some FANG employees
I think this may be something that is unique to a certain size tech company that simply isn't the case at 99.99% of companies with user data in their possession.
Why?
I've worked for a few large tech companies that handled very sensitive customer data, and they didn't allow unsanitized access to it by a significant number of people. Typically (on the dev side, anyway), there was a small designated team (less than 10 people) who were the only ones who had such access. Any dev work that absolutely required access to that data -- which was very rare -- was performed by that team.
I worked at one place that had the development network permanently VPN'd into prod. One day, a developer accidentally configured his local environment to connect to a production queue and database. It was like this for over a week.
A previous company didn't bother with the VPN. They had an AWS environment that predated VPC, so SSH and many other service ports were open to the office IP addresses. And several people's homes, for remote work.
It depends on the company (as with large ones, apparently). I currently work for a small company, and it is no less diligent about this stuff than the major companies I've worked for.
Backups generally have "god mode" access (best description) as they need to backup and restore not just filesystem data, but the audit log data as well.
Most (corp) places I worked, the developers and SysAdmin's working on production servers gave little thought to the backup component apart from making sure the software is installs and runs. ;)
Policy is never useful because even if there is enforcement, there is never 100% perfect enforcement that beats out cryptographic enforcement, at which point policy is no longer needed.
For example Apple can state that your data is end-to-end encrypted and they have no access, and it would be redundant to also have such a policy saying they will not access your data—they can simply say they can't access your data which is a superset of any such policy.
There are policies like "You are not allowed to access user data." and there are policies like, "All access, keystrokes, and applications that have access to user data are logged and those logs are tied to employee IDs. Further the logs are audited and there must by a form 505/2 on file for every access that details the need for the access, what was done with the data, and how the data was handled. If the auditors discover an access in the logs associated with your employee ID and there is no matching 505/2 on file, you will be subject to immediate termination and may be liable in civil and criminal court. Your signature below states that you understand these restrictions, you consent to monitoring of your behavior, and will abide by the policies."
Strong audit trails, logs that cannot changed by being created in an immutable way, logged access at all terminals and entry points. Combined with a separate auditing group that reports through a different chain of command (like through to the general counsel or something) and you have a policy with teeth.
If i publish my crypto wallet private keys and enact a policy that anyone who tries to take the wallet contents will be beaten to death, and get everyone to agree to this policy, then it would be rendered moot when the person who steals the wallet uses it to hire personal body guards.
Through roughly 2000, the principle saving grace was that disk storage was so expensive, and networking so slow, that large quantities of data were unlikely to be found online except in the case of very major organisations. Most financial firms would read data from tape for analysis or marketing programmes, as an example. A major credit card network might have a couple of, say, Sun Starfire class servers onto which a comprehensive union cardholder databset might be assembled and accessed. One friend reported accessing their campus workstation to which a large national medical insurance database was being processed, from the New York Public Library over Telnet (though I believe they didn't actually log in, they did receive the prompt). E-commerce software vendors and systems stored credit card information, which was accessed. Numerous services and datasets fly around all kinds of organisations, with little protection, and were transmitted in unencrypted FTP sessions. Social networks in which NOC addresses were directly accessible from the office network (WiFi access, natch), with millions of members' data directly accessible.
There are many ways to get this wrong. Few to get it right. And most organisations lack the staff, capitalisation, or incentives to do the right thing.
Google are problably among the best. That leaves open the question of how good they are, and what their past practices have been, even in relatively recent years.
Or how they might behave should their advertising monopoly and revenues fail.
I got very lucky here. The work study job in the job placement center was turning job listings that were faxed in into html to post on our fresh new website. Nobody at the time new what HTML was.
A few years earlier I had bought a "Learn HTML in 24 hours book" and I made a Tony Hawk Pro Skater webpage that listed all the special moves. I used a lot if iframes and thought it was pretty good. CSS wasn't really a thing back then. iframes and tables got the job done.
But I got the job and they thought they got very lucky. I worked in this back room with a computer and a fax machine. Jobs listings would be faxed in. I would scan them and let the OCR software try, and then I would clean it up and add some <h> and <b> tags and then do my econ homework for the rest of my shift.
But as the digital stuff become more popular they hired another guy to be in the back room with me. Dude was a bit of a creep and a student came in looking for a part time job that would work around her classes. He kept on going on about how hot she was. A few weeks later he was talking about he signed up for a few of the classes the hot girl was taking.
Every single student record was available on our computers. Names, address, phone numbers, class schedule, SSN, FAFSA data. It was madness.
And I was a lowly fax to html guy.
Agreed, data governance is important for any company to get right, but someone has to have DB access in order to manage it. When you factor in that so many companies derive revenue from the data they generate, then it gets harder.
If you’re Twitter, how can you build the services you need as an architect or data scientist without the actual data?
Very, very few people need root access and the ability to see all raw data. For example, at Google, this is a tiny number of senior SREs and that's it. Your average employees should be using audited UIs running through service accounts that have restricted permissions.
The solution we had wasn’t particularly difficult to set up, and actually made life much easier for everybody, because we’d provided everybody with a very easy to use interface for the data they wanted (much better than the old school shelling into a DB to run your arbitrary SQL statements).
Everything in the data warehouse was anonymized, and people only had access to the schemas they needed (though this was defined quite broadly). Anonymization was handled by our ETL pipeline. When we first set it up, the requirements were pretty simple and we just wrote a little java app to do it. This scaled pretty poorly, and the team ended up putting a proper ETL product in there. I can’t remember which one they used, but there’s a lot of perfectly decent products in that space (even some open source).
I mean, the data is still accessible, it's just not easy to get it willy-nilly without setting off alarms.
So far, I've not yet come across a system where some level of direct admin access isn't needed for at least "last resort" situations. (obviously, only available to a very specific set of trusted people)
I mean, we all know that it's insane to store plaintext passwords. So why is it necessary to store anything as plaintext?
Depending, of course, on how you want to use it.
One could arguably build chat apps and social media that retained no PII. In my opinion, providers retain PII primarily in order to monetize it.
But the problem is that it becomes toxic waste. And it always leaks, eventually. Putting users at risk, and damaging providers' reputations.
Consider the Tox P2P chat app. Each user device runs a Tor onion service. And chats involve only connections among them. Users need disclose no PII. And there's no need for central servers holding PII.
Regarding social media, consider all the "dark markets" that have run as Tor onion services. There's no reason why any sort of social media that you want couldn't be implemented similarly. Although there'd be central servers, there'd be no need for them to handle any PII. Indeed, the fact that "dark markets" handle PII is one of their main weaknesses.
And it's not even necessary to use Tor. One can achieve substantial privacy and anonymity using nested VPN chains, with far less risk of attracting unwanted attention.