HNHacker News
TopNewBestAskShowJobs

throwawaygoog10

49 karma · joined February 16, 2019

Anonymous Google Engineer
submissionscomments
throwawaygoog10··on Google's differential privacy library
You seem to be missing what differential privacy is. It's not about collection of data, it's about the _use_ of that data. It's no secret that Google has an incredible amount of logging data, but the ways we can use it are very limited. Folks seem to be under the impression that we can wily-nily just go ahead and build products that harvest everything about you and link up the dots across organizations. That's so funny, because it'd make things so much easier sometimes. :P

Instead, we have very strict privacy rules and experts to review the designs for the use of this data. If I even want to train a ML model over real data I have to have an approved privacy review that shows how you maintain privacy.

Where I use differential privacy algorithms in my line of work is to do ad-hoc analysis over suggestions placed in front of users. I have dimensions to aggregate across, but I want to ensure that no one bucket can deanonymize a user. k-anonymity used to be the thing (e.g. if a bucket has <50 people in it, that's too few), but even a large bucket can deanonymize users which is where k-anonymity comes in. I sincerely don't care who the users are, I just want to know how our features get used to try and save them more time.

Do I have access to the underlying logs? Yes. Can I use that to make decisions? No. I can however use the anonymized data to make decisions, and even store that longer than the underlying data exists (most logs exist for <14d).

Differential privacy also makes it possible to train models like SmartCompose by ensuring that the tokens it trains over are diffuse enough to not point back to any one person.

> I'd personally bet that differential privacy techniques that actually give users notable information-theoretic anonymity are very rarely used by Google in general.

For existing things, sure. They did their best, but this is new, reified research. As they're replaced they're being replaced by features which use differential privacy techniques.

throwawaygoog10··on Unlimited Google Drive storage by splitting binary files into base64
It isn't truly an oversight, it's an abuse of the fact that Docs/Sheets/Slides are not counted toward your quota. Their storage model is a little more complicated than a standard stream of bytes like an image or a text file.
throwawaygoog10··on Unlimited Google Drive storage by splitting binary files into base64
Should you want to move from Google services, the best way of ensuring you keep your data is to use Takeout [1], which exports your documents as both doc and html files.

[1] https://takeout.google.com

throwawaygoog10··on Unlimited Google Drive storage by splitting binary files into base64
:)
throwawaygoog10··on “Google just started mass banning/limiting Archive Team downloads”
When we turned down the G+ API we injected synthetic errors in increasing fractions of requests a few days ahead of the actual turndown to alert users. I'm not sure if its the same case here, but it's possible that's what's happening.
throwawaygoog10··on Paul Vixie thinks more people should be running their own DNS servers
What exactly is it that you believe is the truth then?
throwawaygoog10··on Paul Vixie thinks more people should be running their own DNS servers
Disclaimer: I work for Google, but not on DNS or Gmail.

> Now, Google does claim they don't track DNS requests. But consider why that is? Once upon a time they didn't scan Gmail content either, but that was before GMail dominated the webmail space.

You seem to assume that it's a singular organization with a unified agenda, but this really isn't the case. It's the same thing about when folks assume Google looks at your Drive files to recommend ads to you -- it isn't true, there's different motives there.

Drive: we want to sell you storage, your data isn't scanned (except for viruses). Google DNS: speed up DNS, which improves load times, which improves the overall web experience. Photos: Ditto, we want to sell you storage.

Performance is a feature, and most ISP resolvers are junk. Worse, many of those resolvers like to inject their own NXDOMAIN pages. :\

You could argue that Google DNS does positively impact Ads, but only in the respect that faster DNS resolution helps ads load faster too. Overall, I see it as one of those "long term greedy" (my own words) strategies.

As a privacy-conscious Googler myself, I've taken a look at Google DNS to convince myself that it's what it says on the tin. As far as I can tell it is, but I don't expect you to take my word for it. What logging exists is extremely temporary (short-term debugging.)

Re: Gmail, this isn't true either. Sure, there's still processing of your emails (we receive your email, scan it for spam), but it isn't used for Gmail ads. The public perception of this was so bad and the incremental improvement in ad quality so low, that now ads just use your general ad profile. No email scanning involved.

> Software stacks, configuration policies, etc will have all evolved to disfavor niche use cases and favor Google, Cloudflare, etc.

This is a different matter entirely, but this isn't _always_ a bad thing. I'm thinking of TCP here, which has almost entirely been ossified by middleboxes. Same for TLS -- TLS development has been hamstrung by these same kinds of middleboxes and "protocol accelerators." This kind of incredible technology position has allowed for the acceleration of HTTP/2 and the development of QUIC (and therefore HTTP/3). Overall, Google has been incredibly open with the development of these and worked to include everyone. I'm sure it's not always that way. Can you bring up some examples where "niche use-cases" have been locked out by Google-driven software stacks and configuration policies?

throwawaygoog10··on Why Google needed to build a graph serving system
I'm sorry your efforts failed internally. Our infrastructure is somewhat ossified these days: the new and exotic are not well accepted. Other than Spanner (which is still working to replace Bigtable), I can't think of a ton of really novel infrastructure that exists now and didn't when you were around. I mean, we don't even do a lot of generic distributed graph processing anymore. Pregel is dead, long live processing everything on a single task with lots of RAM.

I suspect your project would have been really powerful had you gotten the support you needed, but without a (6) or a (7) next to your name it's really hard to convince people of that. I know a number of PAs that would benefit now from structuring their problems in a graph store with arbitrary-depth joins and transactions. I work on one of those and cringe at some of the solutions we've made.

We've forgotten what it's like to need novel solutions to performance problems. Instead, we have services we just throw more RAM at. Ganpati is over 128GB (it might even be more) of RAM now, I suspect a solution like dgraph could solve its problems much more efficiently not to mention scalably.

Good on you for taking your ideas to market. I'm excited to see how your solution evolves.