11,865 karma · joined March 13, 2008
Research: https://www.cs.princeton.edu/~arvindn/
That said, I don't blame people for trying to uncover Satoshi's identity, and it's possible that the techniques in our paper can help. The big caveat, of course, is access to a corpus that includes Satoshi's code labeled with their true identity.
The paper shows that programmers have distinctive styles in their source code that can be extracted to create a fingerprint of coding style. It's not just the obvious stuff like spaces vs. tabs -- parsing the code and looking at the Abstract Syntax Tree is what results in a powerful fingerprint.
We have a follow-up to this paper showing that surprisingly, coding style survives in compiled binaries (even with optimization turned on and debugging symbols removed): https://www.princeton.edu/~aylinc/papers/caliskan-islam_when...
When we released a draft of the Princeton Bitcoin textbook [1], one piece of feedback was that we focused on cryptocurrency technology as it is today, and ignored the juicy and tumultuous history of how the ideas developed over the last few decades. So I invited Jeremy Clark, who's connected to some of this history, to write a preface to the book. If you're interested in the history, you might enjoy that chapter. [2]
Jeremy and I then got together to develop the ideas further, resulting in the present article, where we also provide some commentary on the current blockchain hype and draw lessons for practitioners and academics.
[1] http://bitcoinbook.cs.princeton.edu/
[2] https://d28rh4a8wq0iu5.cloudfront.net/bitcointech/readings/p...
The surprise here isn't that Bitcoin isn't perfectly anonymous. There are two new findings. The first is the extent to which your Bitcoin payment details get leaked to third party trackers. I've been writing about the excesses of third party tracking for years [1], and I'm pretty jaded, but the extent of the leaks surprised me.
The second main finding is that CoinJoin isn't enough to protect yourself. We tested this on our own transactions, but also by coming up with a way to identity essentially all existing CoinJoins on the blockchain and analyzing their anonymity.
Here is the paper http://randomwalker.info/publications/ad-blocking-framework-...
Here is our blog post about it: https://freedom-to-tinker.com/2017/04/14/the-future-of-ad-bl...
Unsurprisingly, today’s statistical machine translation systems reflect existing gender stereotypes. Translations to English from many gender-neutral languages such as Finnish, Estonian, Hungarian, Persian, and Turkish lead to gender-stereotyped sentences. For example, Google Translate converts these Turkish sentences with genderless pronouns: "O bir doktor. O bir hems¸ire." to these English sentences: "He is a doctor. She is a nurse." A test of the 50 occupation words used in the results presented in Figure 1 shows that the pronoun is translated to “he” in the majority of cases and "she" in about a quarter of cases; tellingly, we found that the gender association of the word vectors almost perfectly predicts which pronoun will appear in the translation.
See section on "Effects of bias in NLP applications" http://randomwalker.info/publications/language-bias.pdf
The simplicity and strength of our results suggests a new null hypothesis for explaining origins of prejudicial behavior in humans, namely, the implicit transmission of ingroup/outgroup identity information through language. That is, before providing an explicit or institutional explanation for why individuals make decisions that disadvantage one group with regards to another, one must show that the unjust decision was not a simple outcome of unthinking reproduction of statistical regularities absorbed with language. Similarly, before positing complex models for how prejudicial attitudes perpetuate from one generation to the next or from one group to another, we must check whether simply learning language is sufficient to explain the observed transmission of prejudice. These new null hypotheses are important not because we necessarily expect them to be true in most cases, but because Occam’s razor now requires that we eliminate them, or at least quantify findings about prejudice in comparison to what is explainable from language transmission alone
(The paper has more along these lines.)
Here's the short version:
We view the approach of "debiasing" word embeddings (Bolukbasi et al., 2016) with skepticism. If we view AI as perception followed by action, debiasing alters the AI’s perception (and model) of the world, rather than how it acts on that perception. This gives the AI an incomplete understanding of the world. We see debiasing as "fairness through blindness". It has its place, but also important limits: prejudice can creep back in through proxies (although we should note that Bolukbasi et al. (2016) do consider "indirect bias" in their paper). Efforts to fight prejudice at the level of the initial representation will necessarily hurt meaning and accuracy, and will themselves be hard to adapt as societal understanding of fairness evolves
Direct link to our paper: http://randomwalker.info/publications/language-bias.pdf
But there is one powerful step browsers can take: put stronger privacy protections into private browsing mode, even at the expense of some functionality. Firefox has taken steps in this direction https://blog.mozilla.org/blog/2015/11/03/firefox-now-offers-...
Traditionally all browsers viewed private browsing mode as protecting against local adversaries and not trackers / network adversaries, and in my opinion this was a mistake.
The tool we built to do this research is open-source https://github.com/citp/OpenWPM/ We'd love to work with outside developers to improve it and do new things with it. We've also released the raw data from our study.
As long as they could not see themselves in a mirror, ants with a blue dot painted on their clypeus did not try to remove it. Set in front of a mirror, ants with such a blue dot on their clypeus tried to clean themselves, while ants with a brown painted dot — of the same color as that of their cuticle — on their clypeus and ants with a blue dot on their occiput did not clean themselves. Very young ants did not present such behavior.
And here's our open-source tool that we've been using to do all these privacy measurements: https://github.com/citp/OpenWPM/
We'd love to see pull requests or just other people using our tool for new interesting findings.
* I wish there were a way to fund online education through philanthropy/donations. Coursera being for-profit leaves a bit of a bad taste in the mouth. At a practical level, it complicates what images I can use in my lectures and qualify as fair use.
* After several years the site is far from being at a point where an instructor can log on and upload content. The interface is constantly changing, confusing, and buggy. My university has a dedicated team who help out instructors with putting their material online and even they are often confused about how to edit this or upload that.
Overall I'm glad that Coursera exists and is finding a revenue stream; my own undergraduate education would have been vastly different if I'd had access to the material that's available today.
[1] Bitcoin and Cryptocurrency technologies https://www.coursera.org/course/bitcointech
http://randomwalker.info/teaching/fall-2014-bitcoin/
(P.S. about 200 students enrolled in the first couple of hours. Great to see the level of interest.)
(We collaborated with Mayer on this research.)
Code and data that you can play with to verify these results / do other similar experiments, using our web privacy measurement tool OpenWPM: https://github.com/englehardt/verizon-uidh / https://github.com/citp/OpenWPM
[1] http://weis2013.econinfosec.org/papers/KrollDaveyFeltenWEIS2... (Section 5).
One of our conclusions was that tracking companies switching to HTTPS would help, but a large majority would have to switch to make any difference, because of the sheer number of trackers (Section 4.1). This proposal or something like it is probably necessary if we're to see that magnitude of change.
It studies the same problem as the author here exploited, but goes a lot farther to try and infer things about the targeted individual from the analytics that FB provides to the advertiser. The paper won the Privacy Enhancing Technologies award.
That was back in 2010. If the author succeeded anyway, it seems that Facebook was careless in their implementation of the fix.
On the other hand, add-ons like Ghostery work much better.
[1] https://securehomes.esat.kuleuven.be/~gacar/persistent/index...
[2] https://securehomes.esat.kuleuven.be/~gacar/persistent/the_w...
In fact, the main problem with the paper is the opposite -- the Facebook graphs they use are regional networks borrowed from [1]. This means that Facebook's actual mixing time should be _dramatically higher_ than the times measured in the paper because of the tendency of random walks to get stuck in regional networks. I believe this is the reason their Facebook-A and Facebook-B mixing times are much lower than the others, such as LiveJournal.
Alvisi et al. have a couple of related papers specifically looking at the implications of our new understanding of social-graph random walks for sybil defenses. [2, 3]
[1] https://www.cs.ucsb.edu/~ravenben/publications/pdf/interacti...
[2] http://www.cs.utexas.edu/users/lorenzo/papers/Alvisi13SoK.pd...
[3] http://www.cs.utexas.edu/users/lorenzo/papers/Alvisi14Commun...
Six degrees of separation refers to shortest paths, and fast mixing is decidedly _not_ the intuition behind it.
We can't automate the process (in part because the transformations necessary are much more complex than "scrambling"), but knowledgeable practitioners can go a long way. I'm knowledgeable but not a practitioner, so I'm not the best source.
In #8 of our report (on the Heritage Health data), you'll notice that while I took Khaled El Emam to task for claims about quantifying risk, I do acknowledge that he did a very good job of de-identification. I don't think there's exactly a "canonical reference" (except HIPAA's superficial list of 18 identifiers), but reports written by practitioners like El Emam are probably useful documents.