HNHacker News
TopNewBestAskShowJobs

randomwalker

11,865 karma · joined March 13, 2008

Princeton prof: https://twitter.com/random_walker

Research: https://www.cs.princeton.edu/~arvindn/

submissionscomments
randomwalker··on Bitcoin lecture series (2015)
I'm the lead author of this textbook/lecture series. It's been a couple of years, and we've been thinking about an update. Let me know what topics you'd most like to see. Note that our goal is not to much to teach the details of specific cryptocurrencies as the concepts underlying them. For example, covering Byzantine Fault Tolerance and its application to blockchain protocols is high on my list.
randomwalker··on De-Anonymizing Programmers via Code Stylometry (2015) [pdf]
I've been asked this a lot, especially since I also research and teach cryptocurrency technology [1]. I haven't personally tried to do that. Satoshi clearly wished to maintain their pseudonymity, and I'd rather respect that. I also think Satoshi's pseudonymity is a powerful statement about decentralized cryptocurrencies, namely that their viability rests on their technical merits, with no need to know or trust their creators.

That said, I don't blame people for trying to uncover Satoshi's identity, and it's possible that the techniques in our paper can help. The big caveat, of course, is access to a corpus that includes Satoshi's code labeled with their true identity.

[1] http://randomwalker.info/bitcoin/

randomwalker··on De-Anonymizing Programmers via Code Stylometry (2015) [pdf]
I'm a coauthor of this paper. It was published a couple of years ago; pleasantly surprised to see it here.

The paper shows that programmers have distinctive styles in their source code that can be extracted to create a fingerprint of coding style. It's not just the obvious stuff like spaces vs. tabs -- parsing the code and looking at the Abstract Syntax Tree is what results in a powerful fingerprint.

We have a follow-up to this paper showing that surprisingly, coding style survives in compiled binaries (even with optimization turned on and debugging symbols removed): https://www.princeton.edu/~aylinc/papers/caliskan-islam_when...

randomwalker··on Bitcoin's Academic Pedigree
Coauthor here. Here's some context for how this essay came about.

When we released a draft of the Princeton Bitcoin textbook [1], one piece of feedback was that we focused on cryptocurrency technology as it is today, and ignored the juicy and tumultuous history of how the ideas developed over the last few decades. So I invited Jeremy Clark, who's connected to some of this history, to write a preface to the book. If you're interested in the history, you might enjoy that chapter. [2]

Jeremy and I then got together to develop the ideas further, resulting in the present article, where we also provide some commentary on the current blockchain hype and draw lessons for practitioners and academics.

[1] http://bitcoinbook.cs.princeton.edu/

[2] https://d28rh4a8wq0iu5.cloudfront.net/bitcointech/readings/p...

randomwalker··on Web merchants routinely leak data about Bitcoin purchases
I'm a co-author of this paper. It's available here: https://arxiv.org/pdf/1708.04748.pdf

The surprise here isn't that Bitcoin isn't perfectly anonymous. There are two new findings. The first is the extent to which your Bitcoin payment details get leaked to third party trackers. I've been writing about the excesses of third party tracking for years [1], and I'm pretty jaded, but the extent of the leaks surprised me.

The second main finding is that CoinJoin isn't enough to protect yourself. We tested this on our own transactions, but also by coming up with a way to identity essentially all existing CoinJoins on the blockchain and analyzing their anonymity.

[1] http://randomwalker.info/web-privacy/

randomwalker··on Princeton’s Ad-Blocker May Put an End to the Ad-Blocking Arms Race
Oops! Thanks, fixed.
randomwalker··on Princeton’s Ad-Blocker May Put an End to the Ad-Blocking Arms Race
Coauthor of the paper here. No --- this is not one of the three techniques that we implemented. It was a hypothetical suggestion for future work. Unfortunate that the article didn't make that clear.

Here is the paper http://randomwalker.info/publications/ad-blocking-framework-...

Here is our blog post about it: https://freedom-to-tinker.com/2017/04/14/the-future-of-ad-bl...

randomwalker··on Semantics derived automatically from language corpora contain human-like biases
Coauthor here. Some of the press articles about our work didn't have a lot of nuance (unsurprisingly), but in the paper we're careful about what we say, what we don't say, and what the implications are. Happy to engage in informed discussion :)
randomwalker··on An Empirical Analysis of Linkability in the Monero Blockchain [pdf]
Coauthor here. Someone has been DoSing the paper site(!), so here's a copy for now: https://drive.google.com/file/d/0B59AisMv54waZXRhbE9GV2NDQUE...
randomwalker··on Language contains human biases, and so will machines trained on language corpora
That's definitely one of our main areas for future research. So far, the only part of the paper where we consider other languages is in studying how model bias affects language translation:

Unsurprisingly, today’s statistical machine translation systems reflect existing gender stereotypes. Translations to English from many gender-neutral languages such as Finnish, Estonian, Hungarian, Persian, and Turkish lead to gender-stereotyped sentences. For example, Google Translate converts these Turkish sentences with genderless pronouns: "O bir doktor. O bir hems¸ire." to these English sentences: "He is a doctor. She is a nurse." A test of the 50 occupation words used in the results presented in Figure 1 shows that the pronoun is translated to “he” in the majority of cases and "she" in about a quarter of cases; tellingly, we found that the gender association of the word vectors almost perfectly predicts which pronoun will appear in the translation.

See section on "Effects of bias in NLP applications" http://randomwalker.info/publications/language-bias.pdf

randomwalker··on Language contains human biases, and so will machines trained on language corpora
Coauthor here. The blog post is written in relatively non-technical language for a general audience, but our paper has tons of technical details that HN readers might enjoy. Give it a read!

http://randomwalker.info/publications/language-bias.pdf

randomwalker··on Language contains human biases, and so will machines trained on language corpora
Yes! We address this in the section "Implications for understanding human prejudice".

The simplicity and strength of our results suggests a new null hypothesis for explaining origins of prejudicial behavior in humans, namely, the implicit transmission of ingroup/outgroup identity information through language. That is, before providing an explicit or institutional explanation for why individuals make decisions that disadvantage one group with regards to another, one must show that the unjust decision was not a simple outcome of unthinking reproduction of statistical regularities absorbed with language. Similarly, before positing complex models for how prejudicial attitudes perpetuate from one generation to the next or from one group to another, we must check whether simply learning language is sufficient to explain the observed transmission of prejudice. These new null hypotheses are important not because we necessarily expect them to be true in most cases, but because Occam’s razor now requires that we eliminate them, or at least quantify findings about prejudice in comparison to what is explainable from language transmission alone

(The paper has more along these lines.)

randomwalker··on Language contains human biases, and so will machines trained on language corpora
OP here. We address this argument in detail in our paper, and we're deeply skeptical of it. See the sections titled "Challenges in addressing bias" and "Awareness is better than blindness".

Here's the short version:

We view the approach of "debiasing" word embeddings (Bolukbasi et al., 2016) with skepticism. If we view AI as perception followed by action, debiasing alters the AI’s perception (and model) of the world, rather than how it acts on that perception. This gives the AI an incomplete understanding of the world. We see debiasing as "fairness through blindness". It has its place, but also important limits: prejudice can creep back in through proxies (although we should note that Bolukbasi et al. (2016) do consider "indirect bias" in their paper). Efforts to fight prejudice at the level of the initial representation will necessarily hurt meaning and accuracy, and will themselves be hard to adapt as societal understanding of fairness evolves

Direct link to our paper: http://randomwalker.info/publications/language-bias.pdf

randomwalker··on Online tracking: A 1-million-site measurement and analysis
Personally I think there are so many of these APIs that for the browser to try to prevent the ability to fingerprint is putting the genie back in the bottle.

But there is one powerful step browsers can take: put stronger privacy protections into private browsing mode, even at the expense of some functionality. Firefox has taken steps in this direction https://blog.mozilla.org/blog/2015/11/03/firefox-now-offers-...

Traditionally all browsers viewed private browsing mode as protecting against local adversaries and not trackers / network adversaries, and in my opinion this was a mistake.

randomwalker··on Online tracking: A 1-million-site measurement and analysis
Coauthor here. I lead the research team at Princeton working to uncover online tracking. Happy to answer questions.

The tool we built to do this research is open-source https://github.com/citp/OpenWPM/ We'd love to work with outside developers to improve it and do new things with it. We've also released the raw data from our study.

randomwalker··on Are ants capable of self recognition? [pdf]
The key bit from the abstract:

As long as they could not see themselves in a mirror, ants with a blue dot painted on their clypeus did not try to remove it. Set in front of a mirror, ants with such a blue dot on their clypeus tried to clean themselves, while ants with a brown painted dot — of the same color as that of their cuticle — on their clypeus and ants with a blue dot on their occiput did not clean themselves. Very young ants did not present such behavior.

randomwalker··on Do privacy studies help? A Retrospective look at Canvas Fingerprinting
There's a follow-up to this post here: https://freedom-to-tinker.com/blog/englehardt/the-web-privac...

And here's our open-source tool that we've been using to do all these privacy measurements: https://github.com/citp/OpenWPM/

We'd love to see pull requests or just other people using our tool for new interesting findings.

randomwalker··on As Coursera Evolves, Colleges Stay On and Investors Buy In
I'm the instructor of an upcoming Coursera course [1]. A couple of observations from my point of view:

* I wish there were a way to fund online education through philanthropy/donations. Coursera being for-profit leaves a bit of a bad taste in the mouth. At a practical level, it complicates what images I can use in my lectures and qualify as fair use.

* After several years the site is far from being at a point where an instructor can log on and upload content. The interface is constantly changing, confusing, and buggy. My university has a dedicated team who help out instructors with putting their material online and even they are often confused about how to edit this or upload that.

Overall I'm glad that Coursera exists and is finding a revenue stream; my own undergraduate education would have been vastly different if I'd had access to the material that's available today.

[1] Bitcoin and Cryptocurrency technologies https://www.coursera.org/course/bitcointech

randomwalker··on Princeton Bitcoin and Cryptocurrency Technologies Online Course Now Open
Thanks for the suggestion. We discuss it briefly, and one of the programming assignments is based on Ripple. It could use a deeper discussion, I'll work on that.
randomwalker··on Princeton Bitcoin and Cryptocurrency Technologies Online Course Now Open
Ah, you can't sign up directly on Piazza without a Princeton email, but if you sign up using the Google form link, we'll enroll you on Piazza. (We do this manually; about once a day.)
randomwalker··on Princeton Bitcoin and Cryptocurrency Technologies Online Course Now Open
I'm the lead instructor. Let me know if there are any topics you'd like to see in particular. Here's the syllabus from the version of the class I taught at Princeton:

http://randomwalker.info/teaching/fall-2014-bitcoin/

(P.S. about 200 students enrolled in the first couple of hours. Great to see the level of interest.)

randomwalker··on How Verizon and Turn Defeat Browser Privacy Protections
Some additional research that came out today, with more details of the various things Verizon is doing / plans to do with UIDH: https://freedom-to-tinker.com/blog/englehardt/verizons-track...

(We collaborated with Mayer on this research.)

Code and data that you can play with to verify these results / do other similar experiments, using our web privacy measurement tool OpenWPM: https://github.com/englehardt/verizon-uidh / https://github.com/citp/OpenWPM

randomwalker··on Consensus in Bitcoin: One system, many models
That's a good question. One of the points I'll make in the follow-up post that I promised is that it is indeed possible to capture the motivation of such an attacker in game theory, and in fact, this has been done. [1] However, it makes the model less elegant and introduces parameters. The more of these complexities you wish to model, the less tractable the model becomes.

[1] http://weis2013.econinfosec.org/papers/KrollDaveyFeltenWEIS2... (Section 5).

randomwalker··on Marking HTTP as Non-Secure
One of the surveillance attacks pointed out in the post is the NSA piggybacking on advertising cookies. Details in the Snowden leaks were scant, so we did some research to figure out just how far the NSA could go with this technique. Very far, as it turns out. Here's a blog post with a link to our research paper: https://freedom-to-tinker.com/blog/dreisman/cookies-that-giv...

One of our conclusions was that tracking companies switching to HTTPS would help, but a large majority would have to switch to make any difference, because of the sheer number of trackers (Section 4.1). This proposal or something like it is probably necessary if we're to see that magnitude of change.

randomwalker··on How I Pranked My Roommate with Eerily Targeted Facebook Ads
The reason that Facebook changed their policy to disallow targeting to a very small audience is because of this paper http://repository.cmu.edu/cgi/viewcontent.cgi?article=1066&c...

It studies the same problem as the author here exploited, but goes a lot farther to try and infer things about the targeted individual from the analytics that FB provides to the advertiser. The paper won the Privacy Enhancing Technologies award.

That was back in 2010. If the author succeeded anyway, it seems that Facebook was careless in their implementation of the fix.

randomwalker··on How a new type of “evercookie” tracks you online
Other threads about this paper:

https://news.ycombinator.com/item?id=8064934

https://news.ycombinator.com/item?id=8147376

randomwalker··on The hidden perils of cookie syncing
I'm one of the authors of the paper. One of our findings was that third-party cookie blocking is only marginally effective. See the tables under cookie syncing in our summary [1] or in our full paper [2]. Trackers bypass cookie blocking in a variety of creative ways. We're currently investigating the bypassing mechanisms in more detail.

On the other hand, add-ons like Ghostery work much better.

[1] https://securehomes.esat.kuleuven.be/~gacar/persistent/index...

[2] https://securehomes.esat.kuleuven.be/~gacar/persistent/the_w...

randomwalker··on Friendship Paradox
No. None of the graphs they use are ego-nets. That would be methodologically ridiculous, as you point out.

In fact, the main problem with the paper is the opposite -- the Facebook graphs they use are regional networks borrowed from [1]. This means that Facebook's actual mixing time should be _dramatically higher_ than the times measured in the paper because of the tendency of random walks to get stuck in regional networks. I believe this is the reason their Facebook-A and Facebook-B mixing times are much lower than the others, such as LiveJournal.

Alvisi et al. have a couple of related papers specifically looking at the implications of our new understanding of social-graph random walks for sybil defenses. [2, 3]

[1] https://www.cs.ucsb.edu/~ravenben/publications/pdf/interacti...

[2] http://www.cs.utexas.edu/users/lorenzo/papers/Alvisi13SoK.pd...

[3] http://www.cs.utexas.edu/users/lorenzo/papers/Alvisi14Commun...

randomwalker··on Friendship Paradox
Social graphs are not fast mixing. This used to be widely assumed but recent empirical measurements have refuted the assumption. [1] If I recall correctly, random walks tend to get stuck in cities and other highly-dense subgraphs.

Six degrees of separation refers to shortest paths, and fast mixing is decidedly _not_ the intuition behind it.

[1] http://syssec.kaist.ac.kr/~yongdaek/doc/imc2010.pdf

randomwalker··on No silver bullet: De-identification still doesn’t work
Regardless of whether PHI can be anonymized in a fool-proof way, we can agree that careful anonymization is better than a superficial one, and so your question is valid and important.

We can't automate the process (in part because the transformations necessary are much more complex than "scrambling"), but knowledgeable practitioners can go a long way. I'm knowledgeable but not a practitioner, so I'm not the best source.

In #8 of our report (on the Heritage Health data), you'll notice that while I took Khaled El Emam to task for claims about quantifying risk, I do acknowledge that he did a very good job of de-identification. I don't think there's exactly a "canonical reference" (except HIPAA's superficial list of 18 identifiers), but reports written by practitioners like El Emam are probably useful documents.

← PreviousPage 2 of 15Next →